Ross ROSS = Recommend OSS · open-source software intelligence for agents

abus-aikorea/voice-pro

Gradio WebUI for creators and developers, featuring key TTS (Edge-TTS, kokoro) and zero-shot Voice Cloning (E2 & F5-TTS, CosyVoice), with Whisper audio processing, YouTube download, Demucs vocal isolation, and multilingual translation. observed · 2026-08-28

github.com/abus-aikorea/voice-pro · homepage · Python · GPL-3.0 (copyleft) observed · 2026-08-28

Health v2 · maintenance only

84/100

  • Activity 92
  • Release rhythm 92
  • Longevity 54
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 14.5
  • age_days: 765
  • days_rel: 52
  • days_push: 52
  • n_releases_24m: 11

Full methodology

Adoption not part of the score

12647 stars · 1832 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Voice-Pro is a Gradio-based web UI for AI speech processing, combining TTS engines (Edge-TTS, kokoro), zero-shot voice cloning (E2/F5-TTS, CosyVoice), Whisper-based transcription, and multilingual translation. It also bundles YouTube downloading, Demucs/MDX-Net vocal isolation, and subtitle generation for dubbing workflows.

Use cases

  • clone a voice from a short audio sample
  • transcribe youtube videos to subtitles
  • dub a video into another language
  • convert text to speech for audiobooks
  • separate vocals from a song for karaoke
  • generate translated subtitles for podcasts
  • create multilingual voiceovers

When to choose

  • you want an all-in-one local WebUI for TTS, voice cloning, transcription, and dubbing
  • you need subtitle generation plus translation in one workflow
  • you want an open-source alternative to ElevenLabs for voice cloning

When to avoid

  • you need a headless API or library to embed in your own code
  • you have no GPU and need fast, large-scale processing
  • you need a production-grade hosted service rather than a local tool

Facets

application · maturity active

speech-recognition tts nlp audio-processing video-processing http-server gui speech-processing media artificial-intelligence windows python self-hosted voice-cloning text-to-speech whisper gradio subtitles dubbing yt-dlp vocal-isolation demucs translation audiobook podcast natural-language-processing localization docker gpu

4 sources

Member repositories

RepositoryRoleHealth v2
abus-aikorea/voice-promain84

For agents

markdown · JSON · MCP: product_card(name="abus-aikorea/voice-pro")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem