Ross ROSS = Recommend OSS · open-source software intelligence for agents

Teradata/kylo

Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologies such as Teradata, Apache Spark and/or Hadoop. Kylo is licensed under Apache 2.0. Contributed by Teradata Inc. observed · 2026-08-28

github.com/Teradata/kylo · homepage · Java · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 3493
  • days_rel: n/a
  • days_push: 1329
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1112 stars · 559 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Kylo is an open-source enterprise data lake management platform for self-service data ingest and preparation, with integrated metadata management, governance, and security. It runs on big data engines such as Hadoop, Apache Spark, and NiFi, and is developed in Java.

Use cases

  • self-service data ingest into a Hadoop data lake
  • build governed ETL pipelines without writing code
  • manage metadata, lineage, and data catalog for a data lake
  • prepare and wrangle data with visual SQL transformations
  • profile and validate data quality during ingestion
  • enforce security and governance best practices on big data platforms

When to choose

  • you run Hadoop/Spark/NiFi-based data lakes and need ingest plus governance in one platform
  • you want business users to self-service ingest and prepare data with a guided UI
  • you need integrated metadata repository, lineage, and data profiling

When to avoid

  • you use modern cloud-native stacks (Snowflake, Databricks, dbt) rather than Hadoop-era infrastructure
  • you need a lightweight pipeline tool rather than a full data lake management platform
  • you require active community development - releases and activity have slowed

Facets

application · maturity maintenance

etl workflow-automation search-engine auth authorization web-framework big-data databases analytics self-hosted jvm self-hosted data-lake metadata-management data-governance data-ingest data-preparation nifi spark hadoop lineage data-profiling data-engineering web-server docker

3 sources

Member repositories

RepositoryRoleHealth v2
Teradata/kylomain32

For agents

markdown · JSON · MCP: product_card(name="Teradata/kylo")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem