Ross ROSS = Recommend OSS · open-source software intelligence for agents

MoonshotAI/Kimi-Audio

Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation observed · 2026-08-28

github.com/MoonshotAI/Kimi-Audio · Python observed · 2026-08-28

Health v2 · maintenance only

31/100

  • Activity 27
  • Release rhythm 35
  • Longevity 35

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 495
  • days_rel: n/a
  • days_push: 438
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

4731 stars · 375 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Kimi-Audio is an open-source audio foundation model (7B parameters) that unifies audio understanding, generation, and speech conversation in a single framework. The repository provides inference code, pretrained and instruct model weights, finetuning examples, and an evaluation toolkit.

Use cases

  • transcribe speech to text with a single model
  • build a voice chat assistant that listens and responds
  • answer questions about audio clips
  • classify emotions or sound events from audio
  • caption audio recordings automatically
  • finetune an audio LLM on custom data

When to choose

  • you need one model covering ASR, audio QA, captioning, emotion recognition, and speech conversation
  • you want state-of-the-art open weights for audio understanding benchmarks
  • you want to finetune a large audio foundation model

When to avoid

  • you need a lightweight production ASR with low latency on CPU
  • you need a permissively licensed model - the repo has no license file
  • you only need simple text-to-speech without understanding

Facets

library · maturity active

speech-recognition audio-processing machine-learning llm-inference tts speech-processing artificial-intelligence large-language-models python audio-foundation-model speech-conversation asr audio-understanding 7b-model moonshot-ai audio gpu linux

1 source

Member repositories

RepositoryRoleHealth v2
MoonshotAI/Kimi-Audiomain31

For agents

markdown · JSON · MCP: product_card(name="MoonshotAI/Kimi-Audio")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem