datajuicer/data-juicer
Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷 observed · 2026-08-28
Health v2 · maintenance only
94/100
- Activity 99
- Release rhythm 96
- Longevity 80
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 19.5
- age_days: 1128
- days_rel: 26
- days_push: 7
- n_releases_24m: 25
Adoption not part of the score
6938 stars · 407 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Data-Juicer is a Python library and data processing system for cleaning, deduplicating, synthesizing, and analyzing data for foundation model training. It provides 200+ composable operators that scale from a laptop to thousand-node clusters.
Use cases
- clean and deduplicate web-scale pre-training corpora for LLMs
- filter and curate instruction-tuning datasets
- prepare domain-specific RAG index data
- synthesize training data for foundation models
- process multimodal datasets for AI training
- analyze and visualize dataset quality before training
- curate agent interaction traces for training
When to choose
- you need to clean, filter, or deduplicate large datasets for LLM pre-training or fine-tuning
- you want composable data processing operators with recipes for foundation model data
- you need data processing that scales from laptop to large clusters
- you are preparing multimodal or synthetic training data
When to avoid
- you need a simple one-off ETL job unrelated to AI/ML data
- you require a fully managed GUI data platform rather than a code-first library
- your data volumes are small and a few pandas scripts would suffice
Facets
library · maturity active
etl data-science data-visualization machine-learning llm-training rag nlp large-language-models data-science artificial-intelligence python cloud cross-platform data-processing foundation-models synthetic-data data-cleaning deduplication multimodal data-pipeline pre-training-data data-engineering docker
2 sources
- readme: https://github.com/datajuicer/data-juicer · fetched 2026-08-28 · f27276109fe6
- homepage: https://datajuicer.github.io/data-juicer/ · fetched 2026-08-29 · fc6c7e5d5afd
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| datajuicer/data-juicer | main | 94 |
For agents
markdown · JSON · MCP: product_card(name="datajuicer/data-juicer")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem