lucidrains/naturalspeech2-pytorch
Implementation of Natural Speech 2, Zero-shot Speech and Singing Synthesizer, in Pytorch observed · 2026-08-28
Health v2 · maintenance only
20/100
- Activity 0
- Release rhythm 8
- Longevity 88
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1232
- days_rel: n/a
- days_push: 1074
- n_releases_24m: 0
Adoption not part of the score
1333 stars · 104 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A PyTorch implementation of NaturalSpeech 2, a zero-shot text-to-speech and singing synthesizer that combines a neural audio codec with a latent diffusion model for non-autoregressive generation. It is a research codebase (marked WIP) published as a pip-installable library for training and sampling speech diffusion models.
Use cases
- train a zero-shot text-to-speech model in pytorch
- synthesize singing voices from text
- generate speech audio with a latent diffusion model
- experiment with neural audio codecs like encodec for TTS
- clone a voice from a short speech prompt
- research non-autoregressive speech synthesis
When to choose
- you want to experiment with or extend the NaturalSpeech 2 architecture in PyTorch
- you need a research-grade zero-shot TTS/singing synthesis codebase
- you want to train a latent-diffusion speech model on your own audio data
When to avoid
- you need a production-ready TTS system with pretrained voices
- you want a simple API to convert text to speech without training
- you need guaranteed stability or long-term support for a WIP research repo
Facets
library · maturity experimental
machine-learning deep-learning tts audio-processing artificial-intelligence deep-learning speech-processing python text-to-speech latent-diffusion zero-shot speech-synthesis singing-synthesis neural-audio-codec pytorch research-implementation natural-language-processing gpu
2 sources
- readme: https://github.com/lucidrains/naturalspeech2-pytorch · fetched 2026-08-28 · 7b891ca2fd8a
- registry_pypi: https://pypi.org/pypi/naturalspeech2-pytorch/json · fetched 2026-08-29 · 1810774c8426
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| lucidrains/naturalspeech2-pytorch | main | 20 |
For agents
markdown · JSON · MCP: product_card(name="lucidrains/naturalspeech2-pytorch")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem