Ross ROSS = Recommend OSS · open-source software intelligence for agents

ytsaurus/ytsaurus

YTsaurus is a scalable and fault-tolerant open-source big data platform. observed · 2026-08-28

github.com/ytsaurus/ytsaurus · homepage · C++ · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

94/100

  • Activity 99
  • Release rhythm 86
  • Longevity 97
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 4
  • age_days: 1367
  • days_rel: 19
  • days_push: 7
  • n_releases_24m: 102

Full methodology

Adoption not part of the score

2200 stars · 218 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

YTsaurus is an open-source, fault-tolerant big data platform combining distributed storage, MapReduce processing, a SQL query engine, and a NoSQL key-value store. It integrates ClickHouse (CHYT) for fast analytics and Apache Spark (SPYT) for ETL, and scales to exabytes of data and millions of CPU cores.

Use cases

  • run MapReduce batch jobs over large datasets
  • run ad hoc SQL analytics without exporting data to an external OLAP system
  • store and query data with a low-latency transactional key-value store
  • build ETL pipelines with Apache Spark or SQL
  • manage GPU clusters for training large machine learning models
  • store transactional metadata for distributed systems
  • deploy a multi-tenant big data platform on Kubernetes

When to choose

  • you need a scalable, fault-tolerant platform combining storage, batch processing, and analytics in one system
  • you want multi-tenant big data infrastructure serving many users on shared hardware
  • you need exabyte-scale storage with no single point of failure
  • you want ClickHouse-style SQL analytics directly on your data lake
  • you need a transactional key-value store for OLTP workloads alongside batch processing

When to avoid

  • you only need a simple single-node database or analytics engine
  • you want a fully managed cloud service without operating your own cluster
  • your workloads are small enough for Postgres, ClickHouse, or Spark alone
  • you need a lightweight embedded or edge deployment

Facets

service · maturity active

database search-engine etl streaming scheduling caching big-data databases microservices analytics self-hosted cloud big-data mapreduce distributed-file-system nosql olap sql-query-engine clickhouse apache-spark key-value-store data-lakehouse data-engineering linux docker kubernetes

4 sources

Member repositories

RepositoryRoleHealth v2
ytsaurus/ytsaurusmain94

For agents

markdown · JSON · MCP: product_card(name="ytsaurus/ytsaurus")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem