# ictnlp/StreamSpeech

StreamSpeech is an “All in One” seamless model for offline and simultaneous speech recognition, speech translation and speech synthesis.

Repository: https://github.com/ictnlp/StreamSpeech
Canonical: https://ross.abutalabs.com/products/streamspeech
Homepage: https://ictnlp.github.io/StreamSpeech-site/
Language: Python
License: MIT
License Family: permissive
Topics: seamless, simultaneous-translation, speech, speech-recognition, speech-synthesis, speech-to-text, speech-translation, translation, all-in-one, machine-translation, streaming-audio, text-to-speech, asr, tts, voice, text-to-audio, non-autoregressive, speech-enhancement, audio-processing, speech-processing
Last push: 2025-06-29T02:06:27+00:00

## Health v2 (maintenance only)
Score: 37/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 29, release rhythm 35, longevity 58
- inputs: {"age_days": 820, "days_push": 431, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1287, forks 105 (observed 2026-08-28T04:04:15.026312+00:00)

## What it is
StreamSpeech is an 'All in One' seamless model for offline and simultaneous speech recognition, speech translation, and speech synthesis, presented in an ACL 2024 paper. It performs streaming ASR, simultaneous speech-to-text and speech-to-speech translation with a single multi-task model, showing intermediate results in real time.

## Use cases
- translate speech to speech in real time
- streaming speech recognition while audio is being spoken
- simultaneous speech-to-text translation
- build a low-latency voice interpreter
- research on simultaneous speech-to-speech translation models
- show live ASR and translation transcripts during a call

## When to choose
- you need state-of-the-art simultaneous (streaming) speech-to-speech translation
- you want one model handling ASR, translation, and synthesis together
- you are doing research on low-latency speech translation

## When to avoid
- you need a production-ready hosted translation API
- you only need simple offline batch transcription with mature tooling
- you lack GPU resources for large speech models

## Facets
- artifact type: library
- maturity: active
- function: speech-recognition, tts, nlp, machine-learning, audio-processing
- domain: speech-processing, machine-learning, artificial-intelligence
- platform: python
- tags: simultaneous-translation, speech-to-speech-translation, streaming, asr, tts, multi-task-learning, research-code, acl-2024, natural-language-processing, linux, gpu

## Member repositories
- ictnlp/StreamSpeech (main) score 37

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:15.026312+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:55:52.068418+00:00, confidence not recorded.
  - readme: https://github.com/ictnlp/StreamSpeech (fetched 2026-08-28T04:04:15.026312+00:00, sha eeee89241e18)
  - homepage: https://ictnlp.github.io/StreamSpeech-site/ (fetched 2026-08-29T12:12:02.929922+00:00, sha 6b0a44d57852)
- Data as of 2026-08-30T08:39:29.467469+00:00.
