# QwenLM/Qwen-MM-Plugins

Make any agent harness multimodal-native.

Repository: https://github.com/QwenLM/Qwen-MM-Plugins
Canonical: https://ross.abutalabs.com/products/qwen-mm-plugins
Language: HTML
License: Apache-2.0
License Family: permissive
Last push: 2026-08-25T07:38:04+00:00

## Health v2 (maintenance only)
Score: 57/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 99, release rhythm 35, longevity 2
- inputs: {"age_days": 35, "days_push": 8, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, young
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2777, forks 166 (observed 2026-08-28T04:07:21.467153+00:00)

## What it is
A collection of native multimodal plugins (Skills plus optional MCP servers) for Qwen models that make any agent harness multimodal-native. Capabilities include image/video understanding, OCR, web search, long-video memory, video editing, and 3D/CAD tooling, installed independently per harness.

## Use cases
- make my coding agent understand images and video
- add OCR and document reading to an agent harness
- transcribe meeting videos with speaker labels
- build memory for long-video question answering
- let an agent drive Blender or FreeCAD
- add web and reverse-image search to an AI agent

## When to choose
- you use Claude Code, Qwen Code, Codex, Gemini CLI, or a supported harness and want multimodal capabilities
- you want independently installable skills with optional MCP servers
- you need vision, audio, video, or 3D tooling without building integrations yourself

## When to avoid
- you need a standalone multimodal application rather than agent plugins
- your agent harness is unsupported and you cannot do manual setup
- you want capabilities without installing external dependencies like ffmpeg or API keys

## Facets
- artifact type: plugin
- maturity: active
- function: mcp, agent-framework, computer-vision, speech-recognition, video-processing, image-processing, ocr, search-engine, rag
- domain: artificial-intelligence, computer-vision, developer-tools
- platform: windows, cli, python
- tags: multimodal, qwen, skills, mcp-server, agent-plugins, video-memory, cad, blender, ai-agents, video, linux, macos

## Member repositories
- QwenLM/Qwen-MM-Plugins (main) score 57

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:21.467153+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T08:16:30.929931+00:00, confidence not recorded.
  - readme: https://github.com/QwenLM/Qwen-MM-Plugins (fetched 2026-08-28T04:07:21.467153+00:00, sha 4b9d90d64252)
- Data as of 2026-08-30T08:39:29.467469+00:00.
