# open-mmlab/Amphion

Amphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.

Repository: https://github.com/open-mmlab/Amphion
Canonical: https://ross.abutalabs.com/products/amphion
Homepage: https://openhlt.github.io/amphion/
Language: Python
License: MIT
License Family: permissive
Topics: audio-generation, audio-synthesis, audioldm, music-generation, naturalspeech2, singing-voice-conversion, speech-synthesis, text-to-audio, text-to-speech, vall-e, voice-conversion, audit, fastspeech2, vits, emilia, maskgct, vocoder
Last push: 2026-03-25T14:11:57+00:00

## Health v2 (maintenance only)
Score: 51/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 74, release rhythm 8, longevity 73
- inputs: {"age_days": 1022, "days_push": 161, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 10271, forks 847 (observed 2026-08-28T04:10:41.419761+00:00)

## What it is
Amphion is an open-source Python toolkit for audio, music, and speech generation, supporting tasks like text-to-speech, voice conversion, singing voice conversion, and text-to-audio. It emphasizes reproducible research and includes visualizations of classic model architectures to help junior researchers learn the field.

## Use cases
- generate speech from text with tts models
- convert one voice to another
- convert singing voice between singers
- generate audio or music from text descriptions
- train or evaluate vocoders
- reproduce research on speech generation models
- learn how tts and audio generation architectures work

## When to choose
- you need a reproducible research platform for speech or audio generation
- you want implementations of models like VALL-E, NaturalSpeech2, FastSpeech2, or VITS
- you are a junior researcher learning audio generation with visualized architectures
- you need vocoders and evaluation metrics bundled with generation models

## When to avoid
- you need a production-ready, low-latency tts service for end users
- you only want a simple pretrained model API without training pipelines
- you work outside audio, music, or speech domains

## Facets
- artifact type: library
- maturity: active
- function: tts, audio-processing, machine-learning, deep-learning, speech-recognition, llm-training
- domain: speech-processing, machine-learning, deep-learning
- platform: python
- tags: text-to-speech, voice-conversion, singing-voice-conversion, text-to-audio, text-to-music, vocoder, audio-generation, research-toolkit, speech-synthesis, model-visualization, audio, natural-language-processing, linux, gpu

## Member repositories
- open-mmlab/Amphion (main) score 51

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:41.419761+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:19:21.886581+00:00, confidence not recorded.
  - readme: https://github.com/open-mmlab/Amphion (fetched 2026-08-28T04:10:41.419761+00:00, sha 0cef2b4020b7)
- Data as of 2026-08-30T08:39:29.467469+00:00.
