Ross ROSS = Recommend OSS · open-source software intelligence for agents

juanmc2005/diart

A python package to build AI-powered real-time audio applications observed · 2026-08-28

github.com/juanmc2005/diart · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

62/100

  • Activity 88
  • Release rhythm 8
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1850
  • days_rel: 567
  • days_push: 75
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

2022 stars · 165 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Diart is a Python framework for building AI-powered real-time audio applications, best known for state-of-the-art streaming speaker diarization. It combines speaker segmentation and embedding models with incremental clustering, and supports custom pipelines, hyper-parameter tuning, benchmarking, and serving over websockets.

Use cases

  • identify who is speaking in a live audio stream
  • real-time speaker diarization for meetings or calls
  • stream voice activity detection over websockets
  • build custom real-time speech AI pipelines
  • benchmark and tune online diarization hyper-parameters
  • live captioning with speaker labels

When to choose

  • you need low-latency, streaming speaker diarization in Python
  • you want to prototype real-time audio AI pipelines with pre-trained models
  • you need to serve audio AI models over websockets

When to avoid

  • you only need offline (batch) diarization of recorded files
  • you need production transcription, which is still marked as coming soon
  • you work outside the Python/PyTorch ecosystem

Facets

framework · maturity active

audio-processing speech-recognition machine-learning deep-learning streaming benchmarking speech-processing machine-learning artificial-intelligence python cross-platform speaker-diarization streaming-audio voice-activity-detection speaker-embedding websockets real-time-speech audio real-time

2 sources

Member repositories

RepositoryRoleHealth v2
juanmc2005/diartmain62

For agents

markdown · JSON · MCP: product_card(name="juanmc2005/diart")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem