Ross ROSS = Recommend OSS · open-source software intelligence for agents

nyrahealth/CrisperWhisper

Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps. observed · 2026-08-28

github.com/nyrahealth/CrisperWhisper · homepage · Python · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

89/100

  • Activity 99
  • Release rhythm 94
  • Longevity 59

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 2
  • age_days: 831
  • days_rel: 41
  • days_push: 10
  • n_releases_24m: 2

Full methodology

Adoption not part of the score

1349 stars · 86 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

CrisperWhisper 2.0 is a controllable speech recognition model and Python library that transcribes audio either verbatim (including fillers, repetitions, stutters, and vocal sounds) or in a cleaned 'intended' form, with precise word-level timestamps. It builds on Whisper with a fast CTranslate2 runtime, multilingual support, and seamless longform transcription.

Use cases

  • transcribe audio verbatim with fillers and stutters included
  • generate clean readable transcripts from speech
  • get word-level timestamps for subtitles or alignment
  • convert clean transcript corpora into verbatim datasets for TTS training
  • analyze disfluencies in clinical speech or stuttering research
  • transcribe long recordings without chunk-boundary artifacts

When to choose

  • you need verbatim transcription that captures disfluencies, fillers, and vocal events
  • precise word-level timing matters, e.g. for subtitles, labeling, or speech analysis
  • you want a controllable choice between verbatim and cleaned output in one model
  • you need multilingual speech recognition with disfluency detection

When to avoid

  • you only need lightweight real-time dictation with minimal compute
  • you require a permissively licensed model for commercial redistribution without checking the custom license
  • you need speaker diarization or full-duplex conversation analytics out of the box

Facets

library · maturity active

speech-recognition nlp machine-learning llm-inference speech-processing machine-learning healthcare python cross-platform asr whisper verbatim-transcription word-level-timestamps disfluency-detection cttranslate2 filler-detection stutter-detection multilingual natural-language-processing gpu

4 sources

Member repositories

RepositoryRoleHealth v2
nyrahealth/CrisperWhispermain89

For agents

markdown · JSON · MCP: product_card(name="nyrahealth/CrisperWhisper")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem