Henry-23/VideoChat
实时交互数字人,可自定义形象与音色,支持音色克隆,对话延迟低至3s。Real-time voice interactive digital human, customizable appearance and voice, supporting voice cloning, with initial package delay as low as 3s. observed · 2026-08-28
Health v2 · maintenance only
48/100
- Activity 57
- Release rhythm 35
- Longevity 48
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 684
- days_rel: n/a
- days_push: 258
- n_releases_24m: 0
Adoption not part of the score
1303 stars · 172 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A real-time voice-interactive digital human application that combines ASR, LLM, TTS, and talking-head generation (MuseTalk) into a low-latency conversational avatar. It supports customizable appearance and voice, including voice cloning, with first-packet latency as low as 3 seconds.
Use cases
- build a real-time talking avatar that responds to voice
- create a custom digital human with my own face and voice
- clone a voice for an interactive virtual assistant
- stream a lip-synced talking head from an LLM chatbot
- run an end-to-end voice-to-avatar pipeline locally
- demo a low-latency conversational digital human
When to choose
- you need a self-hosted real-time digital human with lip sync
- you want voice cloning and customizable avatar appearance
- you have a GPU (8GB+ for cascade, 20GB+ for end-to-end MLLM) and want low first-packet latency
- you want to swap in different ASR, LLM, or TTS components
When to avoid
- you have no GPU or limited VRAM
- you need a production-grade hosted service rather than a demo app
- you need non-Linux deployment support
- you want a polished no-code product instead of a Python/Gradio pipeline
Facets
application · maturity active
speech-recognition tts chatbot llm-inference agent-framework video-processing audio-processing machine-learning artificial-intelligence large-language-models speech-processing computer-vision chatbots deep-learning python self-hosted digital-human talking-head lip-sync voice-cloning real-time musetalk gradio asr thg multimodal linux gpu
1 source
- readme: https://github.com/Henry-23/VideoChat · fetched 2026-08-28 · ac1d32349940
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| Henry-23/VideoChat | main | 48 |
For agents
markdown · JSON · MCP: product_card(name="Henry-23/VideoChat")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem