zai-org/CogVLM2
GPT4V-level open-source multi-modal model based on Llama3-8B observed · 2026-08-28
Health v2 · maintenance only
28/100
- Activity 9
- Release rhythm 35
- Longevity 60
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 845
- days_rel: n/a
- days_push: 548
- n_releases_24m: 0
Adoption not part of the score
2433 stars · 162 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
CogVLM2 is an open-source multi-modal vision-language model family built on Meta-Llama-3-8B-Instruct, offering image and video understanding with up to 8K context and 1344x1344 resolution. The repository provides model weights, inference code, and demos for tasks like VQA, document understanding, and video question answering.
Use cases
- answer questions about images with a vision-language model
- understand and summarize videos with an open-source model
- run a GPT-4V alternative locally on my own GPU
- extract text and answer questions from document images
- build a multimodal chatbot that sees images
- quantize a large vision-language model to fit in 16GB VRAM
When to choose
- you need a self-hosted open-source alternative to GPT-4V for image or video understanding
- you want strong benchmark performance on TextVQA and DocVQA
- you need bilingual (Chinese/English) multimodal understanding
- you want an Apache-2.0 licensed VLM you can fine-tune or deploy
When to avoid
- you need a lightweight model for CPU-only or edge devices
- you want a managed API without managing GPU infrastructure
- you need a general-purpose text-only LLM
- you require the newest multimodal models with broader ecosystem support
Facets
library · maturity maintenance
machine-learning deep-learning llm-inference computer-vision nlp large-language-models computer-vision artificial-intelligence python multimodal vision-language-model video-understanding llama3 image-captioning vqa video gpu linux
1 source
- readme: https://github.com/zai-org/CogVLM2 · fetched 2026-08-28 · 4f72524c3703
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| zai-org/CogVLM2 | main | 28 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem