# MoonInTheRiver/DiffSinger

DiffSinger: Singing Voice Synthesis via Shallow Diffusion Mechanism (SVS & TTS); AAAI 2022; Official code

Repository: https://github.com/MoonInTheRiver/DiffSinger
Canonical: https://ross.abutalabs.com/products/diffsinger
Language: Python
License: MIT
License Family: permissive
Topics: text-to-speech, diffusion-speedup, tts, aaai2022, singing-synthesis, diffusion-model, speech-synthesis, singing-voice-synthesis, singing-voice, singing-voice-database, midi
Last push: 2026-07-24T07:22:44+00:00

## Health v2 (maintenance only)
Score: 65/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 94, release rhythm 8, longevity 100
- inputs: {"age_days": 1720, "days_push": 40, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4851, forks 828 (observed 2026-08-28T04:09:01.427548+00:00)

## What it is
Official PyTorch implementation of DiffSinger, an AAAI 2022 paper on singing voice synthesis and text-to-speech using a shallow diffusion mechanism. It includes training and inference code for SVS (with MIDI input) and TTS (DiffSpeech), plus acceleration plugins like PNDM.

## Use cases
- synthesize singing voice from MIDI and lyrics
- generate speech from text with diffusion models
- train a singing voice synthesis model on the PopCS or Opencpop datasets
- accelerate diffusion-based speech synthesis with PNDM
- reproduce AAAI 2022 DiffSinger paper results
- build a demo of neural singing synthesis

## When to choose
- you need a research-grade implementation of diffusion-based singing voice synthesis
- you want to experiment with shallow diffusion for speech or singing generation
- you need MIDI-conditioned singing synthesis with training code
- you want to reproduce or extend the DiffSinger/DiffSpeech papers

## When to avoid
- you need a production-ready, user-friendly TTS product with pretrained voices and simple APIs
- you want real-time low-latency synthesis out of the box
- you need a maintained general-purpose TTS toolkit rather than a research codebase
- you lack a GPU or ML engineering experience

## Facets
- artifact type: library
- maturity: maintenance
- function: machine-learning, deep-learning, audio-processing, speech-recognition, tts
- domain: artificial-intelligence, deep-learning, speech-processing, machine-learning
- platform: python
- tags: diffusion-model, singing-voice-synthesis, text-to-speech, svs, pytorch, research-code, aaai-2022, midi, audio, linux, gpu

## Member repositories
- MoonInTheRiver/DiffSinger (main) score 65

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:09:01.427548+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:18:24.562242+00:00, confidence not recorded.
  - readme: https://github.com/MoonInTheRiver/DiffSinger (fetched 2026-08-28T04:09:01.427548+00:00, sha 20081bdfcf21)
- Data as of 2026-08-30T08:39:29.467469+00:00.
