NVIDIA-NeMo/Curator
Scalable data pre processing and curation toolkit for LLMs observed · 2026-08-28
Health v2 · maintenance only
86/100
- Activity 99
- Release rhythm 83
- Longevity 64
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 64
- age_days: 902
- days_rel: 37
- days_push: 7
- n_releases_24m: 12
Adoption not part of the score
1736 stars · 319 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
NVIDIA NeMo Curator is a scalable, GPU-accelerated toolkit for preprocessing and curating large text, image, video, and audio datasets for AI model training. It provides repeatable pipelines for loading, filtering, deduplicating, and transforming data that can run on a laptop or across multi-node Ray clusters.
Use cases
- deduplicate large web-scale text corpora for LLM training
- filter and quality-score training data for fine-tuning
- build speech and audio datasets with speaker diarization
- curate image datasets with aesthetic and NSFW filtering
- run data preparation pipelines on multi-node Ray or Slurm clusters
- generate synthetic training data with an in-pipeline LLM inference server
- detect languages and classify documents at scale
When to choose
- you need GPU-accelerated processing of very large multi-modal datasets
- you want repeatable, production-grade data curation pipelines for LLM training
- you need semantic or exact deduplication at scale
- you run on HPC clusters (Slurm) or multi-node Ray setups
When to avoid
- you only need small-scale data cleaning on a CPU laptop
- you want a no-code GUI tool rather than a Python library
- your pipeline is simple enough for pandas or plain scripts
Facets
library · maturity active
etl data-science machine-learning llm-training nlp image-processing audio-processing video-processing streaming machine-learning large-language-models developer-tools python cloud data-curation deduplication gpu-accelerated ray synthetic-data-generation fine-tuning data-quality multi-modal data-engineering natural-language-processing linux docker gpu
1 source
- readme: https://github.com/NVIDIA-NeMo/Curator · fetched 2026-08-28 · c65fba70b6ae
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| NVIDIA-NeMo/Curator | main | 86 |
For agents
markdown · JSON · MCP: product_card(name="NVIDIA-NeMo/Curator")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem