Ross ROSS = Recommend OSS · open-source software intelligence for agents

lhotse-speech/lhotse

Tools for handling multimodal data in machine learning projects. observed · 2026-09-01

github.com/lhotse-speech/lhotse · homepage · Python · Apache-2.0 (permissive) observed · 2026-09-01

Health v2 · maintenance only

89/100

  • Activity 100
  • Release rhythm 68
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 41.5
  • age_days: 2322
  • days_rel: 135
  • days_push: 2
  • n_releases_24m: 11

Full methodology

Adoption not part of the score

1149 stars · 278 forks observed · 2026-09-01

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Lhotse is a Python library for flexible, scalable preparation of multimodal (speech, audio, video, image, text) data for machine learning, part of the next-generation Kaldi ecosystem. It provides data preparation recipes, efficient dataloading with bucketing and dataset blending, and sequential I/O formats for distributed training.

Use cases

  • prepare speech corpora for ASR model training
  • build multimodal audio-text training pipelines in PyTorch
  • blend and bucket audio datasets for efficient dataloading
  • shard audio data for distributed multi-node training
  • convert Kaldi data directories to a Python-friendly format
  • deduplicate and randomize training data across GPUs

When to choose

  • you train speech or audio models in PyTorch and need scalable data pipelines
  • you want Python-centric alternatives to Kaldi data preparation
  • you need efficient on-the-fly bucketing or dataset blending for large audio corpora

When to avoid

  • you need a ready-made speech recognition model rather than data tooling
  • your project is not Python/PyTorch based
  • you only need simple one-off audio file conversions

Facets

library · maturity active

machine-learning audio-processing etl data-science sdk speech-processing machine-learning python cross-platform speech-recognition kaldi pytorch dataloader multimodal data-preparation audio-cuts dataset-blending audio data-engineering natural-language-processing

1 source

Member repositories

RepositoryRoleHealth v2
lhotse-speech/lhotsemain89

For agents

markdown · JSON · MCP: product_card(name="lhotse-speech/lhotse")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem