Ross ROSS = Recommend OSS · open-source software intelligence for agents

SesameAILabs/csm

A Conversational Speech Generation Model observed · 2026-08-28

github.com/SesameAILabs/csm · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

30/100

  • Activity 23
  • Release rhythm 35
  • Longevity 39

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 553
  • days_rel: n/a
  • days_push: 463
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

14720 stars · 1478 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

CSM (Conversational Speech Model) is Sesame's speech generation model that produces conversational audio from text and audio context, using a Llama backbone with a Mimi audio decoder. The repository provides Python code for loading the 1B checkpoint and generating speech, including context-aware multi-speaker conversations.

Use cases

  • generate natural-sounding speech from text
  • build a conversational voice assistant
  • create multi-speaker dialogue audio with context prompts
  • clone speaker identity from audio context segments
  • integrate TTS into Python applications via Hugging Face Transformers

When to choose

  • you need context-aware, conversational-quality speech generation
  • you have a CUDA GPU and want a Python API for TTS
  • you want an open-weights 1B speech model with Llama backbone

When to avoid

  • you need production TTS without GPU hardware
  • you need non-English speech or fine-grained voice control beyond context prompting
  • you need a lightweight CPU-only text-to-speech solution

Facets

library · maturity active

tts speech-recognition machine-learning deep-learning llm-inference speech-processing artificial-intelligence large-language-models python windows conversational-speech voice-generation llama-backbone rvq-audio-codes mimi-codec huggingface audio gpu linux macos

1 source

Member repositories

RepositoryRoleHealth v2
SesameAILabs/csmmain30

For agents

markdown · JSON · MCP: product_card(name="SesameAILabs/csm")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem