modelscope/3D-Speaker
A Repository for Single- and Multi-modal Speaker Verification, Speaker Recognition and Speaker Diarization observed · 2026-08-28
Health v2 · maintenance only
56/100
- Activity 56
- Release rhythm 35
- Longevity 91
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1276
- days_rel: n/a
- days_push: 268
- n_releases_24m: 0
Adoption not part of the score
3121 stars · 268 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
3D-Speaker is an open-source Python toolkit for single- and multi-modal speaker verification, speaker recognition, and speaker diarization, with pretrained models hosted on ModelScope. It includes recipes and benchmarks for models like CAM++, ERes2Net, ECAPA-TDNN, and SDPN, plus a large-scale speech corpus for speech representation research.
Use cases
- verify speaker identity from voice recordings
- diarize multi-speaker meeting audio to see who spoke when
- extract speaker embeddings from audio
- train speaker verification models on VoxCeleb or CNCeleb
- identify spoken language from audio
- run multi-modal speaker recognition combining audio and video
When to choose
- you need state-of-the-art speaker verification or diarization with pretrained models
- you want reproducible recipes and benchmarks on VoxCeleb, CNCeleb, or 3D-Speaker datasets
- you work in Python/PyTorch and want ModelScope model integration
When to avoid
- you need general speech-to-text transcription rather than speaker analysis
- you need a production-ready plug-and-play service without training infrastructure
- you work outside Linux/PyTorch environments
Facets
library · maturity active
speech-recognition machine-learning audio-processing benchmarking sdk speech-processing machine-learning deep-learning python speaker-verification speaker-diarization speaker-recognition voice-embeddings pytorch pretrained-models language-identification voxceleb audio linux gpu
1 source
- readme: https://github.com/modelscope/3D-Speaker · fetched 2026-08-28 · fcb41ae4723e
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| modelscope/3D-Speaker | main | 56 |
For agents
markdown · JSON · MCP: product_card(name="modelscope/3D-Speaker")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem