Ross ROSS = Recommend OSS · open-source software intelligence for agents

apache/gobblin

A distributed data integration framework that simplifies common aspects of big data integration such as data ingestion, replication, organization and lifecycle management for both streaming and batch data ecosystems. observed · 2026-08-28

github.com/apache/gobblin · homepage · Java · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

66/100

  • Activity 95
  • Release rhythm 8
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 4293
  • days_rel: n/a
  • days_push: 33
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2270 stars · 750 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Apache Gobblin is a distributed data integration framework for ingesting, replicating, organizing, and managing lifecycle of data across streaming and batch ecosystems. It runs at petabyte scale in production and supports ELT patterns with sources and sinks like Kafka, HDFS, S3, ADLS, and vendor APIs.

Use cases

  • ingest kafka topics into a data lake on hdfs or s3
  • replicate data between hdfs and s3 or adls
  • bulk load a serving store like couchbase from the data lake
  • integrate salesforce api data into a data warehouse
  • enforce data retention and gdpr deletion policies
  • schedule and orchestrate batch and streaming ingestion jobs
  • compact and deduplicate datasets in a data lake

When to choose

  • you need scalable, fault-tolerant data ingestion into a data lake
  • you want incremental processing with state management and atomic publishing
  • you need both streaming and batch ingestion in one framework
  • you operate at large scale and need battle-tested infrastructure

When to avoid

  • you need general-purpose data transformation like Spark or Flink
  • you need a general workflow orchestrator like Airflow or Dagster
  • you need a data storage system
  • your use case is small-scale simple ETL where a lighter tool suffices

Facets

framework · maturity stable

etl streaming workflow-automation scheduling data-science big-data microservices jvm cloud self-hosted data-ingestion data-lake elt kafka hadoop data-replication data-retention gdpr-compliance apache-project data-engineering automation linux docker

3 sources

Member repositories

RepositoryRoleHealth v2
apache/gobblinmain66

For agents

markdown · JSON · MCP: product_card(name="apache/gobblin")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem