# apache/hop

Hop Orchestration Platform

Repository: https://github.com/apache/hop
Canonical: https://ross.abutalabs.com/products/apache-hop
Homepage: https://hop.apache.org/
Language: Java
License: Apache-2.0
License Family: permissive
Topics: java, streaming, hop, apache, data-integration, etl, orchestration, pipeline, workflow
Last push: 2026-08-26T18:03:41+00:00

## Health v2 (maintenance only)
Score: 94/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 85, longevity 100
- inputs: {"age_days": 2535, "days_push": 7, "days_rel": 21, "gap_med": 67.0, "n_releases_24m": 11}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1447, forks 470 (observed 2026-08-28T04:04:45.337123+00:00)

## What it is
Apache Hop is an open-source data and metadata orchestration platform for visually designing and running data integration pipelines and workflows. Pipelines designed in the Hop Gui can run on the native Hop engine locally or remotely, or on Spark, Flink, Google Dataflow, and AWS EMR via Apache Beam.

## Use cases
- build etl pipelines without writing code
- orchestrate data workflows across environments
- migrate from Pentaho Kettle/PDI
- run data pipelines on Spark or Flink
- visually design data transformations with drag and drop
- schedule and manage data integration jobs
- move data between databases and files

## When to choose
- you need a visual, metadata-driven data integration tool
- you want design-once-run-anywhere across local, Spark, Flink, or Beam runtimes
- you are migrating away from Pentaho Data Integration
- you prefer open source with an active Apache community

## When to avoid
- you need a lightweight code-first library embedded in your application
- your use case is simple scripting better served by Python/pandas
- you require a fully managed cloud ETL service
- you need real-time sub-millisecond stream processing

## Facets
- artifact type: application
- maturity: active
- function: etl, workflow-automation, streaming, data-science, gui, cli, plugin-system
- domain: big-data, developer-tools, self-hosted
- platform: cross-platform, jvm, cloud, self-hosted
- tags: data-integration, orchestration, pipelines, visual-development, apache-beam, spark, flink, metadata-driven, kettle-pdi-successor, data-engineering, automation, desktop, docker

## Member repositories
- apache/hop (main) score 94

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:45.337123+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:36:05.997153+00:00, confidence not recorded.
  - readme: https://github.com/apache/hop (fetched 2026-08-28T04:04:45.337123+00:00, sha 69e121b10f79)
  - homepage: https://hop.apache.org/ (fetched 2026-08-29T11:45:52.760368+00:00, sha 066e86a4f64d)
  - site_page: https://hop.apache.org/manual/latest/getting-started (fetched 2026-08-29T11:45:52.769550+00:00, sha 9fcd46380066)
  - site_page: https://hop.apache.org/docs/architecture (fetched 2026-08-29T11:45:52.775378+00:00, sha 7454a56e0f68)
  - site_page: https://hop.apache.org/docs/roadmap (fetched 2026-08-29T11:45:52.778013+00:00, sha b585aa02ac07)
  - site_page: https://hop.apache.org/tech-manual/latest (fetched 2026-08-29T11:45:52.771403+00:00, sha 1d5323e301d9)
  - site_page: https://hop.apache.org/dev-manual/latest (fetched 2026-08-29T11:45:52.773063+00:00, sha a9d1eea09593)
- Data as of 2026-08-30T08:39:29.467469+00:00.
