Ross ROSS = Recommend OSS · open-source software intelligence for agents

zai-org/CogVLM

a state-of-the-art-level open visual language model | 多模态预训练模型 observed · 2026-08-28

github.com/zai-org/CogVLM · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

28/100

  • Activity 0
  • Release rhythm 35
  • Longevity 77

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1081
  • days_rel: n/a
  • days_push: 826
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

6744 stars · 451 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

CogVLM is an open-source visual language model (17B) combining a vision encoder with a pretrained language model for image understanding and multi-turn dialogue, with CogAgent (18B) extending it for high-resolution GUI agent tasks. The repo provides model checkpoints, CLI and web demo inference, OpenAI Vision-compatible serving, and finetuning code.

Use cases

  • answer questions about images with a vision-language model
  • generate captions for images
  • build a GUI automation agent that understands screenshots
  • finetune a multimodal VLM on custom data
  • serve an OpenAI Vision-compatible image chat API
  • run visual question answering on documents and charts

When to choose

  • you need an open-source VLM for image understanding or VQA
  • you want to build a GUI agent that operates from screenshots
  • you need a self-hosted multimodal model with finetuning support

When to avoid

  • you lack GPU hardware (multi-billion-parameter models need significant VRAM)
  • you need the newest multimodal models rather than a 2024-era release
  • you want a lightweight CPU-only image captioning solution

Facets

library · maturity maintenance

machine-learning deep-learning llm-inference computer-vision nlp artificial-intelligence large-language-models computer-vision python visual-language-model multimodal vlm image-understanding gui-agent cogagent pretrained-model vqa image-captioning natural-language-processing gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
zai-org/CogVLMmain28

For agents

markdown · JSON · MCP: product_card(name="zai-org/CogVLM")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem