Ross ROSS = Recommend OSS · open-source software intelligence for agents

Henry-23/VideoChat

实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package delay as low as 3s. observed · 2026-08-28

github.com/Henry-23/VideoChat · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

48/100

  • Activity 57
  • Release rhythm 35
  • Longevity 48

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 684
  • days_rel: n/a
  • days_push: 258
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1303 stars · 172 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-latency conversational avatar. It supports customizable appearance and voice, including voice cloning, with first-packet latency as low as 3 seconds.

Use cases

  • build a real-time talking avatar that responds to voice
  • create a custom digital human with my own face and voice
  • clone a voice for an interactive virtual assistant
  • stream a lip-synced talking head from an LLM chatbot
  • run an end-to-end voice-to-avatar pipeline locally
  • demo a low-latency conversational digital human

When to choose

  • you need a self-hosted real-time digital human with lip sync
  • you want voice cloning and customizable avatar appearance
  • you have a GPU (8GB+ for cascade, 20GB+ for end-to-end MLLM) and want low first-packet latency
  • you want to swap in different ASR, LLM, or TTS components

When to avoid

  • you have no GPU or limited VRAM
  • you need a production-grade hosted service rather than a demo app
  • you need non-Linux deployment support
  • you want a polished no-code product instead of a Python/Gradio pipeline

Facets

application · maturity active

speech-recognition tts chatbot llm-inference agent-framework video-processing audio-processing machine-learning artificial-intelligence large-language-models speech-processing computer-vision chatbots deep-learning python self-hosted digital-human talking-head lip-sync voice-cloning real-time musetalk gradio asr thg multimodal linux gpu

1 source

Member repositories

RepositoryRoleHealth v2
Henry-23/VideoChatmain48

For agents

markdown · JSON · MCP: product_card(name="Henry-23/VideoChat")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem