datazip-inc/olake
OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3. observed · 2026-08-28
Health v2 · maintenance only
84/100
- Activity 99
- Release rhythm 86
- Longevity 49
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 4.5
- age_days: 687
- days_rel: 12
- days_push: 7
- n_releases_24m: 71
Adoption not part of the score
1431 stars · 248 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
OLake Go is a high-performance open-source EL (extract-load) engine written in Go that replicates databases (PostgreSQL, MySQL, MongoDB, Oracle, MSSQL, DB2), Kafka, and S3 into Apache Iceberg tables or Parquet files on object storage. It supports full refresh, incremental sync, and CDC with parallel chunking, schema evolution, and a self-hosted web UI/CLI, plus companion Iceberg table maintenance (OLake Fusion).
Use cases
- replicate postgres to apache iceberg
- cdc replication from mysql to data lakehouse
- ingest mongodb into parquet on s3
- stream kafka topics to iceberg tables
- build a lakehouse without spark or debezium
- self-hosted elt pipeline for real-time analytics
- automate iceberg table compaction and maintenance
When to choose
- you need fast, scalable database-to-Iceberg ingestion with CDC and incremental sync
- you want vendor-lock-in-free lakehouse tables queryable by any engine (Glue, Hive, REST, JDBC catalogs)
- you want a lightweight self-hosted pipeline with no Spark, Flink, Kafka Connect, or Debezium dependency
- you need a web UI plus CLI for managing sync jobs on your own infrastructure
When to avoid
- you need transformations in the pipeline (OLake is EL only, not ELT/ETL with transforms)
- your source is not among the supported databases, Kafka, or S3
- you need CDC from Oracle or DB2, which is not yet supported
- you prefer a fully managed cloud data pipeline service
Facets
application · maturity active
etl streaming database data-science cli gui big-data databases analytics self-hosted self-hosted go cross-platform cloud change-data-capture apache-iceberg parquet lakehouse data-replication cdc kafka s3 elt data-pipeline data-engineering docker kubernetes
9 sources
- readme: https://github.com/datazip-inc/olake · fetched 2026-08-28 · b4d99e594eea
- homepage: https://olake.io · fetched 2026-08-29 · 03d17a2caa6d
- site_page: https://olake.io/docs · fetched 2026-08-29 · 41a8e56e415a
- site_page: https://olake.io/docs/fusion/getting-started/overview · fetched 2026-08-29 · c6d56124e7b8
- site_page: https://olake.io/docs/getting-started/quickstart · fetched 2026-08-29 · 60f834ced153
- site_page: https://olake.io/docs/benchmarks/ingestion · fetched 2026-08-29 · fa8d32007d42
- site_page: https://olake.io/docs/release/ingestion/v0.9.0 · fetched 2026-08-29 · 188361ee25cc
- site_page: https://olake.io/about-us · fetched 2026-08-29 · 17002c1e264f
- site_page: https://olake.io/contact · fetched 2026-08-29 · 2ac2a5121e96
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| datazip-inc/olake | main | 84 |
For agents
markdown · JSON · MCP: product_card(name="datazip-inc/olake")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem