# jim-schwoebel/voice_datasets

🔊 A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).

Repository: https://github.com/jim-schwoebel/voice_datasets
Canonical: https://ross.abutalabs.com/products/voice_datasets
Homepage: https://voicecomputing.org
License Family: other
Topics: voice-dataset, voice-datasets, audio-dataset, audio-datasets, datasets, dataset, voice, data, voice-computing, voice-control, voice-synthesis, voice-commands, voice-assistant, voice-recognition, voice-chat, voice-activity-detection, voice-conversion, noise
Last push: 2024-06-06T10:47:45+00:00

## Health v2 (maintenance only)
Score: 23/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 8, longevity 100
- inputs: {"age_days": 2597, "days_push": 818, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2222, forks 262 (observed 2026-08-28T04:06:27.657882+00:00)

## What it is
A curated list of 95+ open-source datasets for voice and sound computing, covering speech, audio events, and music. It serves as a reference catalog for researchers and developers building voice-related machine learning applications.

## Use cases
- find open-source speech recognition datasets
- locate audio datasets for training voice models
- find datasets for speech emotion recognition
- find voice activity detection training data
- find datasets for voice synthesis and TTS
- discover speaker diarization datasets

## When to choose
- you need to discover publicly available voice or audio datasets for ML projects
- you are researching speech, emotion, or audio event recognition and want a starting catalog
- you want a curated reference rather than searching Kaggle or academic sources individually

## When to avoid
- you need the actual dataset files - this is a list of links, not a hosted dataset
- you need a maintained, license-verified dataset catalog with guaranteed up-to-date links
- you need non-audio machine learning datasets

## Facets
- artifact type: dataset
- maturity: maintenance
- function: speech-recognition, audio-processing, machine-learning, nlp
- domain: speech-processing, machine-learning, artificial-intelligence
- platform: cross-platform
- tags: awesome-list, voice-datasets, speech-datasets, audio-datasets, curated-list, voice-computing, audio

## Member repositories
- jim-schwoebel/voice_datasets (main) score 23

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:27.657882+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:45:45.864582+00:00, confidence not recorded.
  - readme: https://github.com/jim-schwoebel/voice_datasets (fetched 2026-08-28T04:06:27.657882+00:00, sha 3132ddbfbbab)
  - homepage: https://voicecomputing.org (fetched 2026-08-29T10:26:08.390836+00:00, sha 09d07d8f50ce)
- Data as of 2026-08-30T08:39:29.467469+00:00.
