QwenLM/Qwen3-VL
Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud. observed · 2026-08-28
Health v2 · maintenance only
52/100
- Activity 65
- Release rhythm 35
- Longevity 52
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 734
- days_rel: n/a
- days_push: 215
- n_releases_24m: 0
Adoption not part of the score
19847 stars · 1841 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Qwen3-VL is a series of open-weight multimodal vision-language models from Alibaba's Qwen team, available in Dense and MoE architectures with Instruct and Thinking editions. The repository provides model weights, usage examples, cookbooks, and tooling for running the models for image/video understanding, OCR, visual coding, and GUI agent tasks.
Use cases
- understand images and videos with a vision-language model
- extract text from images with multilingual OCR
- build a GUI agent that operates PC or mobile apps visually
- generate HTML/CSS/JS or Draw.io diagrams from screenshots
- do 2D and 3D visual grounding and spatial reasoning
- answer math and STEM questions with multimodal reasoning
- run long-context document and hours-long video analysis
When to choose
- you need an open-weight multimodal LLM you can self-host or fine-tune
- your app combines text, image, and video understanding in one model
- you need strong OCR across many languages and degraded image conditions
- you want visual agent capabilities like GUI element recognition and tool use
When to avoid
- you only need text-only LLM inference without vision
- you lack GPU resources and cannot use an API instead
- you need a small edge model outside the offered size range
Facets
library · maturity active
machine-learning deep-learning llm-inference ocr computer-vision agent-framework large-language-models computer-vision artificial-intelligence python cross-platform vision-language-model multimodal qwen video-understanding visual-grounding moe open-weights natural-language-processing ai-agents gpu
1 source
- readme: https://github.com/QwenLM/Qwen3-VL · fetched 2026-08-28 · 5a25421d32fb
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| QwenLM/Qwen3-VL | main | 52 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem