apache/gobblin
A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems. observed · 2026-08-28
Health v2 · maintenance only
66/100
- Activity 95
- Release rhythm 8
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 4293
- days_rel: n/a
- days_push: 33
- n_releases_24m: 0
Adoption not part of the score
2270 stars · 750 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Apache Gobblin is a distributed data integration framework for ingesting, replicating, organizing, and managing lifecycle of data across streaming and batch ecosystems. It runs at petabyte scale in production and supports ELT patterns with sources and sinks like Kafka, HDFS, S3, ADLS, and vendor APIs.
Use cases
- ingest kafka topics into a data lake on hdfs or s3
- replicate data between hdfs and s3 or adls
- bulk load a serving store like couchbase from the data lake
- integrate salesforce api data into a data warehouse
- enforce data retention and gdpr deletion policies
- schedule and orchestrate batch and streaming ingestion jobs
- compact and deduplicate datasets in a data lake
When to choose
- you need scalable, fault-tolerant data ingestion into a data lake
- you want incremental processing with state management and atomic publishing
- you need both streaming and batch ingestion in one framework
- you operate at large scale and need battle-tested infrastructure
When to avoid
- you need general-purpose data transformation like Spark or Flink
- you need a general workflow orchestrator like Airflow or Dagster
- you need a data storage system
- your use case is small-scale simple ETL where a lighter tool suffices
Facets
framework · maturity stable
etl streaming workflow-automation scheduling data-science big-data microservices jvm cloud self-hosted data-ingestion data-lake elt kafka hadoop data-replication data-retention gdpr-compliance apache-project data-engineering automation linux docker
3 sources
- readme: https://github.com/apache/gobblin · fetched 2026-08-28 · 48d2b89aeff1
- homepage: https://gobblin.apache.org/ · fetched 2026-08-29 · 50154bdb5505
- site_page: https://gobblin.apache.org/docs/developer-guide/HighLevelConsumer · fetched 2026-08-29 · 8b14cd766b74
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| apache/gobblin | main | 66 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem