# Teradata/kylo

Kylo is a data lake management software platform and framework for enabling scalable enterprise-class data lakes on big data technologies such as Teradata, Apache Spark and/or  Hadoop. Kylo is licensed under Apache 2.0. Contributed by Teradata Inc.

Repository: https://github.com/Teradata/kylo
Canonical: https://ross.abutalabs.com/products/kylo
Homepage: http://kylo.io
Language: Java
License: Apache-2.0
License Family: permissive
Topics: spark, nifi, kylo, data-lake, teradata, hadoop
Last push: 2023-01-12T08:25:19+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 3493, "days_push": 1329, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1112, forks 559 (observed 2026-08-28T04:03:37.953372+00:00)

## What it is
Kylo is an open-source enterprise data lake management platform for self-service data ingest and preparation, with integrated metadata management, governance, and security. It runs on big data engines such as Hadoop, Apache Spark, and NiFi, and is developed in Java.

## Use cases
- self-service data ingest into a Hadoop data lake
- build governed ETL pipelines without writing code
- manage metadata, lineage, and data catalog for a data lake
- prepare and wrangle data with visual SQL transformations
- profile and validate data quality during ingestion
- enforce security and governance best practices on big data platforms

## When to choose
- you run Hadoop/Spark/NiFi-based data lakes and need ingest plus governance in one platform
- you want business users to self-service ingest and prepare data with a guided UI
- you need integrated metadata repository, lineage, and data profiling

## When to avoid
- you use modern cloud-native stacks (Snowflake, Databricks, dbt) rather than Hadoop-era infrastructure
- you need a lightweight pipeline tool rather than a full data lake management platform
- you require active community development - releases and activity have slowed

## Facets
- artifact type: application
- maturity: maintenance
- function: etl, workflow-automation, search-engine, auth, authorization, web-framework
- domain: big-data, databases, analytics, self-hosted
- platform: jvm, self-hosted
- tags: data-lake, metadata-management, data-governance, data-ingest, data-preparation, nifi, spark, hadoop, lineage, data-profiling, data-engineering, web-server, docker

## Member repositories
- Teradata/kylo (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:37.953372+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:42:39.139230+00:00, confidence not recorded.
  - readme: https://github.com/Teradata/kylo (fetched 2026-08-28T04:03:37.953372+00:00, sha b1f0d5865884)
  - homepage: http://kylo.io (fetched 2026-08-29T12:46:23.124159+00:00, sha 8574279c6235)
  - site_page: https://kylo.io/quickstart.html (fetched 2026-08-29T12:46:23.133309+00:00, sha f9f5eed24764)
- Data as of 2026-08-30T08:39:29.467469+00:00.
