open-mmlab/Amphion
Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development. observed · 2026-08-28
Health v2 · maintenance only
51/100
- Activity 74
- Release rhythm 8
- Longevity 73
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1022
- days_rel: n/a
- days_push: 161
- n_releases_24m: 0
Adoption not part of the score
10271 stars · 847 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, singing voice conversion, and text-to-audio. It emphasizes reproducible research and includes visualizations of classic model architectures to help junior researchers learn the field.
Use cases
- generate speech from text with tts models
- convert one voice to another
- convert singing voice between singers
- generate audio or music from text descriptions
- train or evaluate vocoders
- reproduce research on speech generation models
- learn how tts and audio generation architectures work
When to choose
- you need a reproducible research platform for speech or audio generation
- you want implementations of models like VALL-E, NaturalSpeech2, FastSpeech2, or VITS
- you are a junior researcher learning audio generation with visualized architectures
- you need vocoders and evaluation metrics bundled with generation models
When to avoid
- you need a production-ready, low-latency tts service for end users
- you only want a simple pretrained model API without training pipelines
- you work outside audio, music, or speech domains
Facets
library · maturity active
tts audio-processing machine-learning deep-learning speech-recognition llm-training speech-processing machine-learning deep-learning python text-to-speech voice-conversion singing-voice-conversion text-to-audio text-to-music vocoder audio-generation research-toolkit speech-synthesis model-visualization audio natural-language-processing linux gpu
1 source
- readme: https://github.com/open-mmlab/Amphion · fetched 2026-08-28 · 0cef2b4020b7
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| open-mmlab/Amphion | main | 51 |
For agents
markdown · JSON · MCP: product_card(name="open-mmlab/Amphion")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem