Ross ROSS = Recommend OSS · open-source software intelligence for agents

QwenLM/Qwen3-Omni

Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time. observed · 2026-08-28

github.com/QwenLM/Qwen3-Omni · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

52/100

  • Activity 78
  • Release rhythm 35
  • Longevity 24

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 346
  • days_rel: n/a
  • days_push: 132
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

3980 stars · 289 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Qwen3-Omni is a natively end-to-end omni-modal large language model from Alibaba Cloud's Qwen team that understands text, images, audio, and video and generates real-time streaming text and speech responses. The repository provides model weights, Transformers and vLLM usage code, cookbooks, and demos for running the model.

Use cases

  • build a voice assistant that listens and talks back in real time
  • analyze video content and ask questions about it
  • transcribe and understand multilingual audio input
  • generate natural speech output from an LLM
  • run multimodal chat with images, audio, and video inputs
  • serve an omni-modal model with vLLM

When to choose

  • you need a single model handling text, audio, image, and video input with speech output
  • you want real-time streaming speech responses from an open-weight LLM
  • you need multilingual omni-modal understanding with Apache-2.0 licensing

When to avoid

  • you only need text-only chat and want a smaller, cheaper model
  • you lack GPU resources for large multimodal inference
  • you need fine-grained control over separate ASR and TTS pipelines

Facets

library · maturity active

llm-inference speech-recognition tts machine-learning nlp audio-processing video-processing image-processing large-language-models artificial-intelligence speech-processing computer-vision python cross-platform omni-modal multimodal-llm qwen alibaba-cloud speech-generation video-understanding transformers vllm foundation-model natural-language-processing audio gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
QwenLM/Qwen3-Omnimain52

For agents

markdown · JSON · MCP: product_card(name="QwenLM/Qwen3-Omni")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem