pachyderm/pachyderm
Data-Centric Pipelines and Data Versioning observed · 2026-08-28
Health v2 · maintenance only
36/100
- Activity 4
- Release rhythm 40
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 19
- age_days: 4381
- days_rel: 595
- days_push: 576
- n_releases_24m: 6
Adoption not part of the score
6308 stars · 579 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Pachyderm is a data-centric pipeline platform that automates data transformations with built-in data versioning and lineage tracking. It runs language-agnostic, parallelized multi-stage pipelines on Kubernetes, using standard object stores with deduplication for storage.
Use cases
- version large datasets and track lineage across pipeline stages
- build automated data pipelines that trigger on data changes
- run parallelized ETL transformations on Kubernetes
- create reproducible CI/CD workflows for data science and ML
- process multi-stage pipelines across any data type on cloud or on-prem
- deduplicate and store pipeline data in object stores
When to choose
- you need immutable data versioning and lineage for complex multi-stage pipelines
- your team runs data workloads on Kubernetes and wants autoscaling and parallelism
- you want language-agnostic pipeline steps and reproducible data transformations
- you need a CI/CD-style trigger system driven by data changes
When to avoid
- you only need simple scheduled batch jobs without data versioning
- you cannot operate a Kubernetes cluster or object storage backend
- you want a lightweight single-machine ETL tool
- you need a fully open-source license without commercial terms
Facets
application · maturity active
etl data-science workflow-automation container-orchestration version-control streaming big-data data-science machine-learning microservices go cloud self-hosted windows data-versioning data-lineage data-pipelines cicd-for-data object-storage parallel-processing reproducibility data-engineering containers kubernetes docker linux macos
1 source
- readme: https://github.com/pachyderm/pachyderm · fetched 2026-08-28 · c1fb19426f11
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| pachyderm/pachyderm | main | 36 |
For agents
markdown · JSON · MCP: product_card(name="pachyderm/pachyderm")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem