BayLing-Models/BayLing-Speech
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level. observed · 2026-08-28
Health v2 · maintenance only
32/100
- Activity 22
- Release rhythm 35
- Longevity 51
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 722
- days_rel: n/a
- days_push: 472
- n_releases_24m: 0
Adoption not part of the score
3146 stars · 225 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
LLaMA-Omni is an end-to-end speech interaction model built on Llama-3.1-8B-Instruct that generates simultaneous text and speech responses from speech instructions with latency as low as 226ms. It combines a pretrained speech encoder, speech adaptor, LLM, and streaming speech decoder, trained on the InstructS2S-200K dataset.
Use cases
- build a voice assistant that talks to an LLM in real time
- generate speech and text responses from spoken instructions
- run a low-latency speech-to-speech conversation model locally
- fine-tune an open-source LLM with speech capabilities
- research speech-language model architectures
- create a hands-free voice chatbot without transcription
When to choose
- you need open-source real-time voice interaction with an LLM
- you want simultaneous text and speech output from speech input
- you need low-latency speech responses without ASR transcription
- you're researching speech-language models on Llama backbones
When to avoid
- you only need text-based chat without audio
- you need a production-ready hosted voice API rather than a research model
- you lack GPU resources for 8B-parameter inference
- you need non-English speech interaction
Facets
library · maturity active
speech-recognition tts llm-inference machine-learning speech-processing large-language-models artificial-intelligence python cli speech-language-model speech-to-speech multimodal llama voice-assistant streaming-speech-decoder low-latency natural-language-processing gpu linux
6 sources
- readme: https://github.com/BayLing-Models/BayLing-Speech · fetched 2026-08-28 · 8e0d161302a7
- homepage: https://arxiv.org/abs/2409.06666 · fetched 2026-08-29 · fbb5aed1e97b
- site_page: https://info.arxiv.org/about/donate.html · fetched 2026-08-29 · cca9c3a11c56
- site_page: https://info.arxiv.org/about/ourmembers.html · fetched 2026-08-29 · 47cbc55ff1de
- site_page: https://info.arxiv.org/about · fetched 2026-08-29 · a1f16f915a9a
- site_page: https://info.arxiv.org/labs/index.html · fetched 2026-08-29 · b14a8d05a0ec
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| BayLing-Models/BayLing-Speech | main | 32 |
For agents
markdown · JSON · MCP: product_card(name="BayLing-Models/BayLing-Speech")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem