Ross ROSS = Recommend OSS · open-source software intelligence for agents

xzf-thu/Mega-ASR

First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐** observed · 2026-09-03

github.com/xzf-thu/Mega-ASR · homepage · Python observed · 2026-09-03

Health v2 · maintenance only

59/100

  • Activity 100
  • Release rhythm 35
  • Longevity 7

Flags: no_releases young no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 109
  • days_rel: n/a
  • days_push: 0
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1136 stars · 74 forks observed · 2026-09-03

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Mega-ASR is a foundation automatic speech recognition model trained on 2.6M samples spanning 7 atomic acoustic conditions and 54 compound real-world scenarios, using A2S-SFT and DG-WGPO reinforcement learning. The repository provides training code, model weights, and an associated dataset and benchmark for robust in-the-wild speech recognition.

Use cases

  • transcribe speech in noisy real-world environments
  • recognize far-field or reverberant audio
  • evaluate ASR robustness across acoustic conditions
  • train a robust speech recognition model
  • benchmark ASR models against challenging audio
  • download open ASR model weights

When to choose

  • your ASR pipeline fails on noisy, far-field, or distorted real-world audio
  • you need a robust open ASR foundation model with weights and training data
  • you want to benchmark speech recognition under compound acoustic scenarios

When to avoid

  • you need a lightweight ASR for clean studio audio
  • you require a permissive license - the repo has no license
  • you need streaming or low-latency on-device transcription

Facets

library · maturity active

speech-recognition machine-learning deep-learning llm-training data-generation speech-processing machine-learning artificial-intelligence python asr robust-speech-recognition acoustic-simulation foundation-model reinforcement-learning dataset model-weights audio gpu linux

2 sources

Member repositories

RepositoryRoleHealth v2
xzf-thu/Mega-ASRmain59

For agents

markdown · JSON · MCP: product_card(name="xzf-thu/Mega-ASR")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem