Ross ROSS = Recommend OSS · open-source software intelligence for agents

debpalash/VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages. observed · 2026-08-28

github.com/debpalash/VoiceStudio · homepage · Python · AGPL-3.0 (copyleft) observed · 2026-08-28

Health v2 · maintenance only

80/100

  • Activity 99
  • Release rhythm 97
  • Longevity 10

Flags: young

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 0.0
  • age_days: 146
  • days_rel: 20
  • days_push: 9
  • n_releases_24m: 35

Full methodology

Adoption not part of the score

11737 stars · 1846 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation across 646 languages. It bundles 16 TTS and 11 ASR engines with support for CUDA, Apple Silicon MPS/MLX, ROCm, and CPU compute, and exposes an OpenAI-compatible local API on port 3900.

Use cases

  • clone a voice from a short audio clip locally
  • dub videos into multiple languages while keeping speaker timing
  • transcribe audio or video into editable text
  • create chaptered audiobooks from scripts or EPUBs
  • generate realistic text-to-speech without an API subscription
  • design new synthetic voices by describing gender, age, and accent
  • dictate text in real time on the desktop
  • run an OpenAI-compatible local speech API for apps

When to choose

  • you need an ElevenLabs-style voice workflow that never sends data to a server
  • you want to switch between many TTS/ASR engines on your own GPU or Apple Silicon
  • you need multilingual dubbing, transcription, or audiobook production in one desktop app
  • you want a local OpenAI-compatible /v1/audio/speech endpoint for your own tools

When to avoid

  • you need a fully stable, production-hardened product — the project is in active beta
  • you lack a capable GPU or Apple Silicon and need fast, high-quality voice generation
  • you need a hosted cloud API — the cloud contract is only a preview with no production endpoint
  • you require a permissive license for closed-source integration — it is AGPL-3.0

Facets

application · maturity active

tts speech-recognition audio-processing nlp llm-inference api-framework mcp speech-processing artificial-intelligence media cross-platform windows python voice-cloning video-dubbing audiobook dictation transcription elevenlabs-alternative local-first openai-compatible-api cuda mlx voice-design multilingual audio localization macos linux desktop tauri gpu docker

4 sources

Member repositories

RepositoryRoleHealth v2
debpalash/VoiceStudiomain80

For agents

markdown · JSON · MCP: product_card(name="debpalash/VoiceStudio")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem