Ross ROSS = Recommend OSS · open-source software intelligence for agents

XiaomiMiMo/MiMo-Audio

MiMo-Audio: Audio Language Models are Few-Shot Learners observed · 2026-08-28

github.com/XiaomiMiMo/MiMo-Audio · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

57/100

  • Activity 88
  • Release rhythm 35
  • Longevity 24

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 349
  • days_rel: n/a
  • days_push: 77
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1077 stars · 107 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Xiaomi's open-source 7B audio language model family (Base and Instruct) plus a 1.2B RVQ audio tokenizer, pretrained on 100M+ hours of audio to enable few-shot generalization across audio tasks. It jointly models text and audio tokens for understanding, speech generation, TTS, voice conversion, and speech editing.

Use cases

  • build an audio understanding model that answers questions about speech and sound
  • text-to-speech and instruct-TTS with natural voices
  • speech continuation for generating talk shows or recitations
  • voice conversion and style transfer with few examples
  • speech-to-text transcription and audio reasoning benchmarks
  • few-shot audio task generalization without task-specific fine-tuning

When to choose

  • you need a single open-source model covering both audio understanding and generation
  • you want SOTA open-source performance on audio benchmarks like MMAU or MMSU
  • you need few-shot generalization to audio tasks without fine-tuning
  • you want an Apache-2.0 licensed audio LLM with released checkpoints and eval suite

When to avoid

  • you need lightweight on-device audio processing rather than a 7B GPU model
  • you only need simple ASR where Whisper-class models suffice
  • you lack GPU resources for large model inference
  • you need production hardening rather than a research codebase

Facets

library · maturity active

machine-learning audio-processing speech-recognition tts llm-inference artificial-intelligence speech-processing large-language-models python audio-language-model few-shot-learning audio-tokenizer rvq speech-generation instruction-tuning 7b-model multimodal audio natural-language-processing gpu linux

2 sources

Member repositories

RepositoryRoleHealth v2
XiaomiMiMo/MiMo-Audiomain57

For agents

markdown · JSON · MCP: product_card(name="XiaomiMiMo/MiMo-Audio")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem