# QwenLM/Qwen2.5-Omni

Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.

Repository: https://github.com/QwenLM/Qwen2.5-Omni
Canonical: https://ross.abutalabs.com/products/qwen25-omni
Language: Jupyter Notebook
License: Apache-2.0
License Family: permissive
Last push: 2025-06-12T11:03:07+00:00

## Health v2 (maintenance only)
Score: 31/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 26, release rhythm 35, longevity 37
- inputs: {"age_days": 530, "days_push": 447, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4074, forks 327 (observed 2026-08-28T04:08:34.253547+00:00)

## What it is
Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates streaming text and natural speech responses. The repository provides model weights, inference code, quantized variants, and cookbooks for running the 7B model.

## Use cases
- build a voice assistant that sees and hears
- understand video content with an LLM
- real-time speech generation from multimodal input
- run a multimodal model locally on GPU
- audio understanding and reasoning benchmark model
- quantized multimodal inference to save VRAM

## When to choose
- you need one model handling text, image, audio, and video input with speech output
- you want state-of-the-art open-source audio understanding
- you need streaming, real-time conversational responses

## When to avoid
- you only need text-only LLM inference
- you lack a GPU or need CPU-only deployment
- you need a small lightweight model for edge devices

## Facets
- artifact type: library
- maturity: active
- function: machine-learning, deep-learning, speech-recognition, tts, llm-inference, transformers, nlp, computer-vision, video-processing, audio-processing
- domain: large-language-models, deep-learning, speech-processing, computer-vision, artificial-intelligence
- platform: python
- tags: multimodal, qwen, speech-generation, video-understanding, model-weights, quantization, transformers, natural-language-processing, gpu, linux, docker

## Member repositories
- QwenLM/Qwen2.5-Omni (main) score 31

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:34.253547+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:23:28.710051+00:00, confidence not recorded.
  - readme: https://github.com/QwenLM/Qwen2.5-Omni (fetched 2026-08-28T04:08:34.253547+00:00, sha 7bc2a1f829fe)
- Data as of 2026-08-30T08:39:29.467469+00:00.
