# k2-fsa/ZipVoice

Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching

Repository: https://github.com/k2-fsa/ZipVoice
Canonical: https://ross.abutalabs.com/products/zipvoice
Homepage: https://zipvoice.github.io/
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2025-12-02T08:58:26+00:00

## Health v2 (maintenance only)
Score: 43/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 55, release rhythm 35, longevity 31
- inputs: {"age_days": 439, "days_push": 274, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1045, forks 143 (observed 2026-08-28T04:03:21.599827+00:00)

## What it is
ZipVoice is a series of fast, high-quality zero-shot text-to-speech models based on flow matching, with a compact 123M-parameter Zipformer-based architecture. It supports voice cloning in Chinese and English, single-speaker and dialogue generation, and includes a distilled variant for faster inference.

## Use cases
- clone a voice from a short audio sample
- generate speech from text in Chinese or English
- generate multi-speaker dialogue audio
- run fast TTS on CPU or GPU with few sampling steps
- build an audiobook or narration pipeline with cloned voices
- generate cross-lingual speech with a speaker prompt

## When to choose
- you need fast, small-footprint zero-shot TTS with high speaker similarity
- you need Chinese and English speech synthesis or dialogue generation
- you want an open-source Apache-2.0 TTS model with checkpoints and inference code

## When to avoid
- you need languages beyond Chinese and English
- you need a managed cloud TTS API rather than self-hosted models
- you need real-time streaming synthesis out of the box

## Facets
- artifact type: library
- maturity: active
- function: tts, machine-learning, deep-learning, speech-recognition
- domain: speech-processing, machine-learning
- platform: python, cross-platform
- tags: zero-shot-tts, voice-cloning, flow-matching, zipformer, speech-synthesis, multilingual, natural-language-processing, gpu

## Member repositories
- k2-fsa/ZipVoice (main) score 43

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:21.599827+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:02:05.751107+00:00, confidence not recorded.
  - readme: https://github.com/k2-fsa/ZipVoice (fetched 2026-08-28T04:03:21.599827+00:00, sha 53030125d877)
  - homepage: https://zipvoice.github.io/ (fetched 2026-08-29T13:03:18.729446+00:00, sha 54966d8e7f30)
- Data as of 2026-08-30T08:39:29.467469+00:00.
