Ross ROSS = Recommend OSS · open-source software intelligence for agents

apache/hadoop

Apache Hadoop observed · 2026-08-28

github.com/apache/hadoop · homepage · Java · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

77/100

  • Activity 99
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 4388
  • days_rel: n/a
  • days_push: 7
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

15640 stars · 9242 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Apache Hadoop is an open-source framework for reliable, scalable distributed storage (HDFS) and processing (MapReduce, YARN) of large data sets across clusters of commodity computers. It scales from single servers to thousands of machines and handles failures at the application layer.

Use cases

  • store and process petabytes of data across a cluster
  • run batch MapReduce jobs on large datasets
  • set up a distributed file system for a data lake
  • schedule and manage cluster compute resources with YARN
  • build an on-premises big data platform
  • run data pipelines over cloud object storage like S3

When to choose

  • you need proven, battle-tested distributed storage and batch processing at scale
  • you want a mature ecosystem with HDFS, YARN, and MapReduce under one project
  • you need fault-tolerant processing on commodity hardware without external orchestration

When to avoid

  • you only need lightweight analytics that Spark, DuckDB, or a cloud warehouse handles more simply
  • you want real-time low-latency stream processing as the primary workload
  • you cannot operate the operational overhead of a JVM-based cluster

Facets

framework · maturity stable

etl streaming file-system database scheduling developer-tools big-data microservices analytics jvm cross-platform cloud hdfs mapreduce yarn distributed-computing data-lake cluster-computing apache data-engineering linux docker

10 sources

Member repositories

RepositoryRoleHealth v2
apache/hadoopmain77

For agents

markdown · JSON · MCP: product_card(name="apache/hadoop")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem