Ross ROSS = Recommend OSS · open-source software intelligence for agents

datajuicer/data-juicer

Data processing for and with foundation models! 🍎 🍋 🌽 ➡️ ➡️🍸 🍹 🍷 observed · 2026-08-28

github.com/datajuicer/data-juicer · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

94/100

  • Activity 99
  • Release rhythm 96
  • Longevity 80
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 19.5
  • age_days: 1128
  • days_rel: 26
  • days_push: 7
  • n_releases_24m: 25

Full methodology

Adoption not part of the score

6938 stars · 407 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Data-Juicer is a Python library and data processing system for cleaning, deduplicating, synthesizing, and analyzing data for foundation model training. It provides 200+ composable operators that scale from a laptop to thousand-node clusters.

Use cases

  • clean and deduplicate web-scale pre-training corpora for LLMs
  • filter and curate instruction-tuning datasets
  • prepare domain-specific RAG index data
  • synthesize training data for foundation models
  • process multimodal datasets for AI training
  • analyze and visualize dataset quality before training
  • curate agent interaction traces for training

When to choose

  • you need to clean, filter, or deduplicate large datasets for LLM pre-training or fine-tuning
  • you want composable data processing operators with recipes for foundation model data
  • you need data processing that scales from laptop to large clusters
  • you are preparing multimodal or synthetic training data

When to avoid

  • you need a simple one-off ETL job unrelated to AI/ML data
  • you require a fully managed GUI data platform rather than a code-first library
  • your data volumes are small and a few pandas scripts would suffice

Facets

library · maturity active

etl data-science data-visualization machine-learning llm-training rag nlp large-language-models data-science artificial-intelligence python cloud cross-platform data-processing foundation-models synthetic-data data-cleaning deduplication multimodal data-pipeline pre-training-data data-engineering docker

2 sources

Member repositories

RepositoryRoleHealth v2
datajuicer/data-juicermain94

For agents

markdown · JSON · MCP: product_card(name="datajuicer/data-juicer")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem