Ross ROSS = Recommend OSS · open-source software intelligence for agents

apache/amoro

Apache Amoro(incubating) is a Lakehouse management system built on open data lake formats. observed · 2026-09-03

github.com/apache/amoro · homepage · Java · Apache-2.0 (permissive) observed · 2026-09-03

Health v2 · maintenance only

73/100

  • Activity 100
  • Release rhythm 23
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 146.5
  • age_days: 1512
  • days_rel: 356
  • days_push: 0
  • n_releases_24m: 3

Full methodology

Adoption not part of the score

1171 stars · 395 forks observed · 2026-09-03

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Apache Amoro (incubating) is a Lakehouse management system built on open data lake formats like Iceberg, Paimon, and Mixed-Hive. It provides a unified catalog service, self-optimizing table maintenance, and management tools that work with Flink, Spark, and Trino compute engines.

Use cases

  • manage iceberg tables across flink spark and trino
  • automatically compact small files in a data lake
  • unified catalog service for multiple compute engines
  • build a lakehouse with stream and batch fusion
  • expire old data and reduce lake storage costs
  • cdc ingestion into lakehouse tables
  • deploy lakehouse management on kubernetes

When to choose

  • you use multiple table formats (Iceberg, Paimon, Mixed-Hive) and want unified management
  • you need automatic table optimization like compaction, sorting, and deduplication
  • you want a single catalog for Flink, Spark, and Trino
  • you need real-time CDC ingestion with millisecond-to-second SLAs via LogStore

When to avoid

  • you only use a single engine with native Iceberg tooling and need no extra management layer
  • you need a fully graduated Apache top-level project with guaranteed stability
  • your stack is outside the JVM/Hadoop ecosystem
  • you want a lightweight embedded library rather than a deployed management service

Facets

service · maturity active

database etl streaming workflow-automation monitoring big-data databases analytics jvm self-hosted cloud lakehouse iceberg paimon hudi table-format catalog-service self-optimizing flink spark trino data-lake apache-incubator data-engineering docker kubernetes

10 sources

Member repositories

RepositoryRoleHealth v2
apache/amoromain73

For agents

markdown · JSON · MCP: product_card(name="apache/amoro")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem