# QwenLM/Qwen3-VL

Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.

Repository: https://github.com/QwenLM/Qwen3-VL
Canonical: https://ross.abutalabs.com/products/qwen3-vl
Language: Jupyter Notebook
License: Apache-2.0
License Family: permissive
Last push: 2026-01-30T04:47:30+00:00

## Health v2 (maintenance only)
Score: 52/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 65, release rhythm 35, longevity 52
- inputs: {"age_days": 734, "days_push": 215, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 19847, forks 1841 (observed 2026-08-28T04:11:29.314782+00:00)

## What it is
Qwen3-VL is a series of open-weight multimodal vision-language models from Alibaba's Qwen team, available in Dense and MoE architectures with Instruct and Thinking editions. The repository provides model weights, usage examples, cookbooks, and tooling for running the models for image/video understanding, OCR, visual coding, and GUI agent tasks.

## Use cases
- understand images and videos with a vision-language model
- extract text from images with multilingual OCR
- build a GUI agent that operates PC or mobile apps visually
- generate HTML/CSS/JS or Draw.io diagrams from screenshots
- do 2D and 3D visual grounding and spatial reasoning
- answer math and STEM questions with multimodal reasoning
- run long-context document and hours-long video analysis

## When to choose
- you need an open-weight multimodal LLM you can self-host or fine-tune
- your app combines text, image, and video understanding in one model
- you need strong OCR across many languages and degraded image conditions
- you want visual agent capabilities like GUI element recognition and tool use

## When to avoid
- you only need text-only LLM inference without vision
- you lack GPU resources and cannot use an API instead
- you need a small edge model outside the offered size range

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, llm-inference, ocr, computer-vision, agent-framework
- domain: large-language-models, computer-vision, artificial-intelligence
- platform: python, cross-platform
- tags: vision-language-model, multimodal, qwen, video-understanding, visual-grounding, moe, open-weights, natural-language-processing, ai-agents, gpu

## Member repositories
- QwenLM/Qwen3-VL (main) score 52

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:11:29.314782+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T16:59:25.347157+00:00, confidence not recorded.
  - readme: https://github.com/QwenLM/Qwen3-VL (fetched 2026-08-28T04:11:29.314782+00:00, sha 5a25421d32fb)
- Data as of 2026-08-30T08:39:29.467469+00:00.
