# abus-aikorea/voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation.

Repository: https://github.com/abus-aikorea/voice-pro
Canonical: https://ross.abutalabs.com/products/voice-pro
Homepage: https://www.wctokyoseoul.com
Language: Python
License: GPL-3.0
License Family: copyleft
Topics: faster-whisper, tts, whisper, gradio, subtitles, transcription, translator, webui, speech-recognition, speech-synthesis, speech-to-text, text-to-speech, yt-dlp, voice-cloning, podcasts, audiobook, voice-conversion, karaoke, whisperx
Last push: 2026-07-13T01:28:10+00:00

## Health v2 (maintenance only)
Score: 84/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 92, release rhythm 92, longevity 54
- inputs: {"age_days": 765, "days_push": 52, "days_rel": 52, "gap_med": 14.5, "n_releases_24m": 11}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 12647, forks 1832 (observed 2026-08-28T04:10:59.500495+00:00)

## What it is
Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, CosyVoice), Whisper-based transcription, and multilingual translation. It also bundles YouTube downloading, Demucs/MDX-Net vocal isolation, and subtitle generation for dubbing workflows.

## Use cases
- clone a voice from a short audio sample
- transcribe youtube videos to subtitles
- dub a video into another language
- convert text to speech for audiobooks
- separate vocals from a song for karaoke
- generate translated subtitles for podcasts
- create multilingual voiceovers

## When to choose
- you want an all-in-one local WebUI for TTS, voice cloning, transcription, and dubbing
- you need subtitle generation plus translation in one workflow
- you want an open-source alternative to ElevenLabs for voice cloning

## When to avoid
- you need a headless API or library to embed in your own code
- you have no GPU and need fast, large-scale processing
- you need a production-grade hosted service rather than a local tool

## Facets
- artifact type: application
- maturity: active
- function: speech-recognition, tts, nlp, audio-processing, video-processing, http-server, gui
- domain: speech-processing, media, artificial-intelligence
- platform: windows, python, self-hosted
- tags: voice-cloning, text-to-speech, whisper, gradio, subtitles, dubbing, yt-dlp, vocal-isolation, demucs, translation, audiobook, podcast, natural-language-processing, localization, docker, gpu

## Member repositories
- abus-aikorea/voice-pro (main) score 84

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:59.500495+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:13:53.702640+00:00, confidence not recorded.
  - readme: https://github.com/abus-aikorea/voice-pro (fetched 2026-08-28T04:10:59.500495+00:00, sha 5003fd8016a8)
  - homepage: https://www.wctokyoseoul.com (fetched 2026-08-29T08:10:19.796363+00:00, sha 3b04d1c1db74)
  - site_page: https://www.wctokyoseoul.com/ko/about (fetched 2026-08-29T08:10:19.807641+00:00, sha 98b137d61955)
  - site_page: https://www.wctokyoseoul.com/ko/faq (fetched 2026-08-29T08:10:19.805848+00:00, sha 2db5c9cdfbb3)
- Data as of 2026-08-30T08:39:29.467469+00:00.
