Ross ROSS = Recommend OSS · open-source software intelligence for agents

datazip-inc/olake

OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3. observed · 2026-08-28

github.com/datazip-inc/olake · homepage · Go · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

84/100

  • Activity 99
  • Release rhythm 86
  • Longevity 49
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 4.5
  • age_days: 687
  • days_rel: 12
  • days_push: 7
  • n_releases_24m: 71

Full methodology

Adoption not part of the score

1431 stars · 248 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

OLake Go is a high-performance open-source EL (extract-load) engine written in Go that replicates databases (PostgreSQL, MySQL, MongoDB, Oracle, MSSQL, DB2), Kafka, and S3 into Apache Iceberg tables or Parquet files on object storage. It supports full refresh, incremental sync, and CDC with parallel chunking, schema evolution, and a self-hosted web UI/CLI, plus companion Iceberg table maintenance (OLake Fusion).

Use cases

  • replicate postgres to apache iceberg
  • cdc replication from mysql to data lakehouse
  • ingest mongodb into parquet on s3
  • stream kafka topics to iceberg tables
  • build a lakehouse without spark or debezium
  • self-hosted elt pipeline for real-time analytics
  • automate iceberg table compaction and maintenance

When to choose

  • you need fast, scalable database-to-Iceberg ingestion with CDC and incremental sync
  • you want vendor-lock-in-free lakehouse tables queryable by any engine (Glue, Hive, REST, JDBC catalogs)
  • you want a lightweight self-hosted pipeline with no Spark, Flink, Kafka Connect, or Debezium dependency
  • you need a web UI plus CLI for managing sync jobs on your own infrastructure

When to avoid

  • you need transformations in the pipeline (OLake is EL only, not ELT/ETL with transforms)
  • your source is not among the supported databases, Kafka, or S3
  • you need CDC from Oracle or DB2, which is not yet supported
  • you prefer a fully managed cloud data pipeline service

Facets

application · maturity active

etl streaming database data-science cli gui big-data databases analytics self-hosted self-hosted go cross-platform cloud change-data-capture apache-iceberg parquet lakehouse data-replication cdc kafka s3 elt data-pipeline data-engineering docker kubernetes

9 sources

Member repositories

RepositoryRoleHealth v2
datazip-inc/olakemain84

For agents

markdown · JSON · MCP: product_card(name="datazip-inc/olake")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem