# apache/gravitino

World's most powerful open data catalog for building a high-performance, geo-distributed and federated metadata lake.

Repository: https://github.com/apache/gravitino
Canonical: https://ross.abutalabs.com/products/gravitino
Homepage: https://gravitino.apache.org
Language: Java
License: Apache-2.0
License Family: permissive
Topics: datalake, lakehouse, metadata, federated-query, stratosphere, metalake, skycomputing, data-catalog, ai-catalog, model-catalog, opendatacatalog
Last push: 2026-08-26T14:06:25+00:00

## Health v2 (maintenance only)
Score: 89/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 78, longevity 87
- inputs: {"age_days": 1229, "days_push": 7, "days_rel": 65, "gap_med": 48.5, "n_releases_24m": 13}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3188, forks 916 (observed 2026-08-28T04:07:48.071142+00:00)

## What it is
Apache Gravitino is a high-performance, geo-distributed, federated metadata lake that provides unified metadata management across diverse data sources such as Hive, MySQL, PostgreSQL, HDFS, and S3. It offers end-to-end data governance including access control, auditing, and discovery, with multi-engine support for query engines like Trino, Spark, and Flink.

## Use cases
- unify metadata across hive mysql and s3 in one catalog
- federated metadata discovery across data lakes and clouds
- manage access control and auditing for all data sources
- query lakehouse tables with trino spark or flink without changing sql
- track ai models and features alongside data assets
- share metadata across regions with a geo-distributed deployment
- replace hive metastore with a modern open data catalog

## When to choose
- you need a single catalog spanning multiple metadata sources, engines, and clouds
- you want direct metadata management where changes reflect immediately in underlying systems
- you need geo-distributed metadata sharing across regions or cloud providers
- you want unified governance like access control and auditing over data and AI assets

## When to avoid
- you only need a simple single-source metastore for one engine
- you require a lightweight embedded catalog with no server component
- your stack is entirely non-JVM and you cannot run a Java 17 service
- you need mature AI asset management today, as it is still work in progress

## Facets
- artifact type: service
- maturity: active
- function: database, search-engine, auth, api-framework, middleware
- domain: databases, big-data, analytics, self-hosted, developer-tools
- platform: jvm, self-hosted, cloud
- tags: data-catalog, metadata-lake, lakehouse, federated-metadata, data-governance, geo-distributed, trino, spark, flink, hive-metastore, iceberg, ai-asset-management, data-engineering, linux, macos, docker

## Member repositories
- apache/gravitino (main) score 89

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:48.071142+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:24:58.353602+00:00, confidence not recorded.
  - readme: https://github.com/apache/gravitino (fetched 2026-08-28T04:07:48.071142+00:00, sha c0ae95c36640)
  - homepage: https://gravitino.apache.org (fetched 2026-08-29T09:38:48.276093+00:00, sha 09fe8bc980f7)
  - site_page: https://gravitino.apache.org/docs/1.3.0 (fetched 2026-08-29T09:38:48.285098+00:00, sha 34ca78487350)
- Data as of 2026-08-30T08:39:29.467469+00:00.
