Ross ROSS = Recommend OSS · open-source software intelligence for agents

ABexit/ASR-LLM-TTS

This is a speech interaction system built on an open-source model, integrating ASR, LLM, and TTS in sequence. The ASR model is SenceVoice, the LLM models are QWen2.5-0.5B/1.5B, and there are three TTS models: CosyVoice, Edge-TTS, and pyttsx3 observed · 2026-08-28

github.com/ABexit/ASR-LLM-TTS · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

60/100

  • Activity 85
  • Release rhythm 35
  • Longevity 47

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 659
  • days_rel: n/a
  • days_push: 91
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1271 stars · 206 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

An open-source Python speech interaction pipeline that chains SenseVoice ASR, Qwen2.5 LLMs, and TTS engines (CosyVoice, Edge-TTS, pyttsx3) for real-time voice conversations. It adds features like wake-word detection, speaker/voiceprint recognition with CAM++, conversation history memory, interruption handling, and multimodal audio/video input via Qwen2-VL.

Use cases

  • build a local voice assistant with speech recognition and tts
  • real-time speech-to-speech conversation with an llm
  • custom wake word detection for a voice assistant
  • speaker identification for voice chat
  • multimodal voice assistant that understands images and video
  • interruptible real-time voice chat on gpu

When to choose

  • you want a fully open-source, locally runnable ASR-LLM-TTS voice pipeline in Python
  • you need Chinese-focused speech interaction with wake words and speaker recognition
  • you want to experiment with different TTS engines or multimodal (audio/video) input

When to avoid

  • you need a production-grade, scalable voice service rather than a research/demo codebase
  • you have no GPU or cannot download models from HuggingFace/ModelScope
  • you need English-first voice assistant tooling with polished packaging

Facets

application · maturity active

speech-recognition tts llm-inference audio-processing chatbot artificial-intelligence speech-processing large-language-models chatbots python cross-platform asr voice-assistant sensevoice qwen cosyvoice edge-tts voice-activity-detection speaker-recognition wake-word real-time gpu

1 source

Member repositories

RepositoryRoleHealth v2
ABexit/ASR-LLM-TTSmain60

For agents

markdown · JSON · MCP: product_card(name="ABexit/ASR-LLM-TTS")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem