# QwenAudio/Fun-ASR

Open-source LLM-based ASR model family for Chinese, dialect, accent, and multilingual speech, with FunASR, vLLM, streaming, and llama.cpp runtimes.

Repository: https://github.com/QwenAudio/Fun-ASR
Canonical: https://ross.abutalabs.com/products/fun-asr
Homepage: https://huggingface.co/spaces/FunAudioLLM/Fun-ASR-Nano
Language: C
License: Apache-2.0
License Family: permissive
Topics: audio-language-model, pytorch, speaker-diarization, speech-recognition, fun-asr, asr, chinese-dialects, multilingual-asr, real-time-asr, speech-to-text, transcription, whisper-alternative, 31-languages, llm-asr, audio, funasr, gguf, llama-cpp, on-device, fun-asr-nano
Last push: 2026-08-19T03:33:31+00:00

## Health v2 (maintenance only)
Score: 81/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 98, release rhythm 94, longevity 18
- inputs: {"age_days": 262, "days_push": 14, "days_rel": 40, "gap_med": 6, "n_releases_24m": 6}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1496, forks 147 (observed 2026-08-28T04:04:53.555509+00:00)

## What it is
Fun-ASR is a family of open-source LLM-based end-to-end speech recognition models from Tongyi Lab, covering Chinese, dialects, accents, and 31 languages via the Fun-ASR-Nano and Fun-ASR-MLT-Nano checkpoints. It integrates with FunASR for inference and serving, and supports vLLM batch inference, streaming, and llama.cpp on-device deployment.

## Use cases
- transcribe Chinese speech with dialect and accent robustness
- convert audio recordings to text in 31 languages
- run real-time streaming speech recognition
- perform speaker diarization on meeting audio
- deploy an on-device Whisper alternative with llama.cpp
- batch transcribe large audio datasets with vLLM

## When to choose
- you need strong Chinese, dialect, or accented speech recognition
- you want multilingual ASR in a single compact 800M model
- you need streaming or real-time transcription
- you want on-device inference via GGUF/llama.cpp

## When to avoid
- you need speaker verification or full speech analytics beyond diarization
- you require languages outside the supported 31
- you need a managed cloud ASR service rather than self-hosted models

## Facets
- artifact type: library
- maturity: active
- function: speech-recognition, machine-learning, llm-inference, audio-processing
- domain: speech-processing, machine-learning
- platform: python, cross-platform, cli
- tags: asr, speech-to-text, transcription, chinese-dialects, multilingual, funasr, vllm, streaming, speaker-diarization, llama-cpp, gguf, on-device, whisper-alternative, natural-language-processing, audio, gpu

## Member repositories
- QwenAudio/Fun-ASR (main) score 81

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:53.555509+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:33:11.833765+00:00, confidence not recorded.
  - readme: https://github.com/QwenAudio/Fun-ASR (fetched 2026-08-28T04:04:53.555509+00:00, sha 2a8b8621df3e)
  - homepage: https://huggingface.co/spaces/FunAudioLLM/Fun-ASR-Nano (fetched 2026-08-29T11:38:23.987579+00:00, sha 04ee629a85ee)
- Data as of 2026-08-30T08:39:29.467469+00:00.
