Ross ROSS = Recommend OSS · open-source software intelligence for agents

innnky/emotional-vits

无需情感标注的情感可控语音合成模型,基于VITS observed · 2026-08-28

github.com/innnky/emotional-vits · Jupyter Notebook · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1493
  • days_rel: n/a
  • days_push: 1252
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1392 stars · 170 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Emotional VITS is a voice synthesis (TTS) model based on VITS that enables emotion-controllable speech generation without requiring manual emotion labels in the training dataset. It extracts emotion embeddings from audio using a pretrained emotion extraction model and feeds them into a modified TextEncoder, allowing inference with a reference audio to control the emotional tone of synthesized speech.

Use cases

  • synthesize speech with controllable emotion from a reference audio clip
  • train an emotional TTS model on a dataset without emotion annotations
  • cluster audio files by emotional similarity to pick reference clips
  • fine-tune an existing VITS checkpoint to add emotion control
  • build multi-speaker TTS with per-speaker emotional expression

When to avoid

  • you need text-driven emotion control using words like 'excited' or 'calm' without providing a reference audio
  • you need a production-ready TTS service with an API out of the box
  • you cannot provide a reference audio clip at inference time
  • you need real-time low-latency synthesis on constrained hardware

Facets

library · maturity maintenance

tts speech-recognition machine-learning deep-learning audio-processing speech-processing machine-learning python cross-platform vits emotional-tts voice-synthesis emotion-embedding speech-synthesis jupyter-notebook audio natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
innnky/emotional-vitsmain32

For agents

markdown · JSON · MCP: product_card(name="innnky/emotional-vits")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem