Ross ROSS = Recommend OSS · open-source software intelligence for agents

QwenLM/Qwen2.5-Omni

Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation. observed · 2026-08-28

github.com/QwenLM/Qwen2.5-Omni · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

31/100

  • Activity 26
  • Release rhythm 35
  • Longevity 37

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 530
  • days_rel: n/a
  • days_push: 447
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

4074 stars · 327 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Qwen2.5-Omni is an end-to-end multimodal model from Alibaba's Qwen team that understands text, images, audio, and video, and generates streaming text and natural speech responses. The repository provides model weights, inference code, quantized variants, and cookbooks for running the 7B model.

Use cases

  • build a voice assistant that sees and hears
  • understand video content with an LLM
  • real-time speech generation from multimodal input
  • run a multimodal model locally on GPU
  • audio understanding and reasoning benchmark model
  • quantized multimodal inference to save VRAM

When to choose

  • you need one model handling text, image, audio, and video input with speech output
  • you want state-of-the-art open-source audio understanding
  • you need streaming, real-time conversational responses

When to avoid

  • you only need text-only LLM inference
  • you lack a GPU or need CPU-only deployment
  • you need a small lightweight model for edge devices

Facets

library · maturity active

machine-learning deep-learning speech-recognition tts llm-inference transformers nlp computer-vision video-processing audio-processing large-language-models deep-learning speech-processing computer-vision artificial-intelligence python multimodal qwen speech-generation video-understanding model-weights quantization transformers natural-language-processing gpu linux docker

1 source

Member repositories

RepositoryRoleHealth v2
QwenLM/Qwen2.5-Omnimain31

For agents

markdown · JSON · MCP: product_card(name="QwenLM/Qwen2.5-Omni")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem