Ross ROSS = Recommend OSS · open-source software intelligence for agents

jianchang512/clone-voice

A sound cloning tool with a web interface, using your voice or any sound to record audio / 一个带web界面的声音克隆工具,使用你的音色或任意声音来录制音频 observed · 2026-08-28

github.com/jianchang512/clone-voice · homepage · Python · NOASSERTION (other) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 39
  • Release rhythm 8
  • Longevity 72

Flags: archived no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1018
  • days_rel: n/a
  • days_push: 370
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

8989 stars · 984 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A voice cloning tool with a web interface built on the coqui.ai xtts_v2 model, letting users synthesize speech in any voice from text or convert one voice into another. It supports 16 languages, microphone recording, and optional CUDA acceleration, with a precompiled Windows build for easy use.

Use cases

  • clone my voice to read text aloud
  • convert text to speech using a specific voice sample
  • change one audio recording to sound like another voice
  • record my voice from the browser and use it for TTS
  • generate dubbed audio in 16 languages from a script
  • convert srt subtitle text into spoken audio

When to choose

  • you want a simple point-and-click voice cloning tool without GPU requirements
  • you need multilingual text-to-speech with a custom or cloned voice
  • you want to run voice conversion locally with a precompiled Windows binary
  • you need to record a 5-20 second voice sample and synthesize speech from it

When to avoid

  • you need production-grade Chinese synthesis quality (English works best, Chinese is mediocre)
  • you require a permissively licensed model - xtts_v2 uses the Coqui Public Model License
  • you cannot provide a stable proxy for downloading large models from Hugging Face
  • you need real-time streaming TTS or API-first integration rather than a web UI

Facets

application · maturity active

tts speech-recognition audio-processing machine-learning gui speech-processing artificial-intelligence media windows python cross-platform voice-cloning xtts-v2 speech-synthesis voice-conversion web-ui cuda-acceleration multilingual audio linux macos gpu

2 sources

Member repositories

RepositoryRoleHealth v2
jianchang512/clone-voicemain10

For agents

markdown · JSON · MCP: product_card(name="jianchang512/clone-voice")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem