lhotse-speech/lhotse
Tools for handling multimodal data in machine learning projects. observed · 2026-09-01
Health v2 · maintenance only
89/100
- Activity 100
- Release rhythm 68
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 41.5
- age_days: 2322
- days_rel: 135
- days_push: 2
- n_releases_24m: 11
Adoption not part of the score
1149 stars · 278 forks observed · 2026-09-01
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Lhotse is a Python library for flexible, scalable preparation of multimodal (speech, audio, video, image, text) data for machine learning, part of the next-generation Kaldi ecosystem. It provides data preparation recipes, efficient dataloading with bucketing and dataset blending, and sequential I/O formats for distributed training.
Use cases
- prepare speech corpora for ASR model training
- build multimodal audio-text training pipelines in PyTorch
- blend and bucket audio datasets for efficient dataloading
- shard audio data for distributed multi-node training
- convert Kaldi data directories to a Python-friendly format
- deduplicate and randomize training data across GPUs
When to choose
- you train speech or audio models in PyTorch and need scalable data pipelines
- you want Python-centric alternatives to Kaldi data preparation
- you need efficient on-the-fly bucketing or dataset blending for large audio corpora
When to avoid
- you need a ready-made speech recognition model rather than data tooling
- your project is not Python/PyTorch based
- you only need simple one-off audio file conversions
Facets
library · maturity active
machine-learning audio-processing etl data-science sdk speech-processing machine-learning python cross-platform speech-recognition kaldi pytorch dataloader multimodal data-preparation audio-cuts dataset-blending audio data-engineering natural-language-processing
1 source
- readme: https://github.com/lhotse-speech/lhotse · fetched 2026-09-01 · 668c3eb16da8
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| lhotse-speech/lhotse | main | 89 |
For agents
markdown · JSON · MCP: product_card(name="lhotse-speech/lhotse")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem