# debpalash/VoiceStudio

VoiceStudio is the open-source, fully-local ElevenLabs alternative — voice cloning, voice design, video dubbing, dictation, transcription & audiobook creation in 646 languages.

Repository: https://github.com/debpalash/VoiceStudio
Canonical: https://ross.abutalabs.com/products/voicestudio
Homepage: https://voicestudio.sh
Language: Python
License: AGPL-3.0
License Family: copyleft
Topics: tts, voice-cloning, voice-generation, voice-ai, ai, cuda, mlx, huggingface, transcription, translate, workflow, omnivoice-studio, dubbing, audiobook, elevenlabs-alternative, local-first, speech-to-text, tauri, text-to-speech, voicestudio
Last push: 2026-08-24T12:41:40+00:00

## Health v2 (maintenance only)
Score: 80/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 97, longevity 10
- inputs: {"age_days": 146, "days_push": 9, "days_rel": 20, "gap_med": 0.0, "n_releases_24m": 35}
- flags: young
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 11737, forks 1846 (observed 2026-08-28T04:10:49.512920+00:00)

## What it is
VoiceStudio is an open-source, fully-local desktop application (built with Tauri and Python) that provides voice cloning, voice design, video dubbing, dictation, transcription, and audiobook creation across 646 languages. It bundles 16 TTS and 11 ASR engines with support for CUDA, Apple Silicon MPS/MLX, ROCm, and CPU compute, and exposes an OpenAI-compatible local API on port 3900.

## Use cases
- clone a voice from a short audio clip locally
- dub videos into multiple languages while keeping speaker timing
- transcribe audio or video into editable text
- create chaptered audiobooks from scripts or EPUBs
- generate realistic text-to-speech without an API subscription
- design new synthetic voices by describing gender, age, and accent
- dictate text in real time on the desktop
- run an OpenAI-compatible local speech API for apps

## When to choose
- you need an ElevenLabs-style voice workflow that never sends data to a server
- you want to switch between many TTS/ASR engines on your own GPU or Apple Silicon
- you need multilingual dubbing, transcription, or audiobook production in one desktop app
- you want a local OpenAI-compatible /v1/audio/speech endpoint for your own tools

## When to avoid
- you need a fully stable, production-hardened product — the project is in active beta
- you lack a capable GPU or Apple Silicon and need fast, high-quality voice generation
- you need a hosted cloud API — the cloud contract is only a preview with no production endpoint
- you require a permissive license for closed-source integration — it is AGPL-3.0

## Facets
- artifact type: application
- maturity: active
- function: tts, speech-recognition, audio-processing, nlp, llm-inference, api-framework, mcp
- domain: speech-processing, artificial-intelligence, media, cross-platform
- platform: windows, python
- tags: voice-cloning, video-dubbing, audiobook, dictation, transcription, elevenlabs-alternative, local-first, openai-compatible-api, cuda, mlx, voice-design, multilingual, audio, localization, macos, linux, desktop, tauri, gpu, docker

## Member repositories
- debpalash/VoiceStudio (main) score 80

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:49.512920+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:15:21.803474+00:00, confidence not recorded.
  - readme: https://github.com/debpalash/VoiceStudio (fetched 2026-08-28T04:10:49.512920+00:00, sha e872f686a8bc)
  - homepage: https://voicestudio.sh (fetched 2026-08-29T08:13:32.041433+00:00, sha 42a60b6d1477)
  - site_page: https://voicestudio.sh/docs (fetched 2026-08-29T08:13:32.050879+00:00, sha c60c5ef48570)
  - site_page: https://voicestudio.sh/docs/quickstart (fetched 2026-08-29T08:13:32.052736+00:00, sha b0fabd050795)
- Data as of 2026-08-30T08:39:29.467469+00:00.
