NVIDIA-NeMo/Speech
A scalable generative AI framework built for researchers and developers working on Large Language Models, Multimodal, and Speech AI (Automatic Speech Recognition and Text-to-Speech) observed · 2026-08-28
Health v2 · maintenance only
98/100
- Activity 99
- Release rhythm 96
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 19
- age_days: 2585
- days_rel: 27
- days_push: 7
- n_releases_24m: 24
Adoption not part of the score
18337 stars · 3595 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
NVIDIA NeMo Speech is an open-source Python framework for building, training, and deploying speech, audio, and multimodal language models, covering ASR, TTS, speaker tasks, and speech-aware LLMs. It ships production-ready pretrained checkpoints, modular neural components, and scalable multi-GPU/multi-node training via PyTorch Lightning with Hydra-based YAML configuration.
Use cases
- transcribe audio to text with streaming asr
- convert text to natural speech with tts
- diarize who spoke when in multi-speaker audio
- fine-tune speech recognition models on custom data
- build speech-aware large language models
- enhance and separate audio signals
- train tts models on multiple languages
- run speaker recognition and verification
When to choose
- you need state-of-the-art pretrained ASR or TTS models with a path to production
- you want scalable multi-GPU/multi-node training for speech models
- you need streaming ASR with controllable latency
- you are researching speech LLMs or multimodal audio models
- you want speaker diarization, recognition, and verification in one toolkit
When to avoid
- you only need simple audio playback or editing rather than AI models
- you have no GPU and need fast inference (CPU-only is slow)
- you want a lightweight single-purpose ASR library without a large dependency stack
- you need non-PyTorch frameworks like TensorFlow or JAX
Facets
framework · maturity active
speech-recognition tts machine-learning deep-learning audio-processing llm-training sdk speech-processing machine-learning deep-learning artificial-intelligence python cloud asr tts speaker-diarization speech-translation pytorch-lightning nvidia pretrained-models speech-llm voice-activity-detection forced-alignment audio natural-language-processing gpu linux docker
10 sources
- readme: https://github.com/NVIDIA-NeMo/Speech · fetched 2026-08-28 · 7542fca65f19
- homepage: https://docs.nvidia.com/nemo/speech/nightly/index.html · fetched 2026-08-29 · 88940b7bbfb3
- site_page: https://docs.nvidia.com/nemo/speech/nightly/starthere/install.html · fetched 2026-08-29 · 01e754bc554d
- site_page: https://docs.nvidia.com/nemo/speech/nightly/features/parallelisms.html · fetched 2026-08-29 · ffe16f442542
- site_page: https://docs.nvidia.com/nemo/speech/nightly/features/mixed_precision.html · fetched 2026-08-29 · a27ec8c772f6
- site_page: https://docs.nvidia.com/nemo/speech/nightly/asr/speaker_diarization/resources.html · fetched 2026-08-29 · b5e462cc79f8
- site_page: https://docs.nvidia.com/nemo/speech/nightly/asr/speaker_recognition/resources.html · fetched 2026-08-29 · fe734b54483c
- site_page: https://docs.nvidia.com/nemo/speech/nightly/asr/ssl/resources.html · fetched 2026-08-29 · b8ad5cdef481
- site_page: https://docs.nvidia.com/nemo/speech/nightly/asr/speech_classification/resources.html · fetched 2026-08-29 · 42abd7fde7af
- site_page: https://docs.nvidia.com/nemo/speech/nightly/tts/intro.html · fetched 2026-08-29 · 3338cc790bf9
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| NVIDIA-NeMo/Speech | main | 98 |
For agents
markdown · JSON · MCP: product_card(name="NVIDIA-NeMo/Speech")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem