Ross ROSS = Recommend OSS · open-source software intelligence for agents

zai-org/CogVLM2

GPT4V-level open-source multi-modal model based on Llama3-8B observed · 2026-08-28

github.com/zai-org/CogVLM2 · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 9
  • Release rhythm 35
  • Longevity 60

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 845
  • days_rel: n/a
  • days_push: 548
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2433 stars · 162 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

CogVLM2 is an open-source multi-modal vision-language model family built on Meta-Llama-3-8B-Instruct, offering image and video understanding with up to 8K context and 1344x1344 resolution. The repository provides model weights, inference code, and demos for tasks like VQA, document understanding, and video question answering.

Use cases

  • answer questions about images with a vision-language model
  • understand and summarize videos with an open-source model
  • run a GPT-4V alternative locally on my own GPU
  • extract text and answer questions from document images
  • build a multimodal chatbot that sees images
  • quantize a large vision-language model to fit in 16GB VRAM

When to choose

  • you need a self-hosted open-source alternative to GPT-4V for image or video understanding
  • you want strong benchmark performance on TextVQA and DocVQA
  • you need bilingual (Chinese/English) multimodal understanding
  • you want an Apache-2.0 licensed VLM you can fine-tune or deploy

When to avoid

  • you need a lightweight model for CPU-only or edge devices
  • you want a managed API without managing GPU infrastructure
  • you need a general-purpose text-only LLM
  • you require the newest multimodal models with broader ecosystem support

Facets

library · maturity maintenance

machine-learning deep-learning llm-inference computer-vision nlp large-language-models computer-vision artificial-intelligence python multimodal vision-language-model video-understanding llama3 image-captioning vqa video gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
zai-org/CogVLM2main28

For agents

markdown · JSON · MCP: product_card(name="zai-org/CogVLM2")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem