apache/seatunnel
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool. observed · 2026-08-28
Health v2 · maintenance only
82/100
- Activity 99
- Release rhythm 50
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 97
- age_days: 3315
- days_rel: 172
- days_push: 7
- n_releases_24m: 6
Adoption not part of the score
9587 stars · 2375 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Apache SeaTunnel is a distributed, high-performance data integration platform for synchronizing massive amounts of data across hundreds of sources and sinks, supporting batch, streaming, CDC, and multimodal workloads. It runs on its own Zeta engine or on Flink and Spark, with an EtLT model, exactly-once guarantees, and automatic schema evolution.
Use cases
- sync data from mysql to clickhouse
- real-time cdc replication from postgres to a data warehouse
- batch load files from s3 into a data lake
- stream kafka events into elasticsearch
- full database synchronization across hundreds of tables
- ingest multimodal data like images and text into a vector database
- migrate data between oltp databases and olap warehouses
When to choose
- you need to move large volumes of data between many heterogeneous systems with one tool
- you need CDC, batch, and streaming synchronization with exactly-once guarantees
- you want schema evolution handled automatically without pausing pipelines
- you want to run the same connectors on Zeta, Flink, or Spark
- you need resource-efficient real-time sync of many small tables
When to avoid
- you need heavy in-flight transformations rather than lightweight EtLT transforms
- you only have a single simple source-to-sink copy that a database tool handles natively
- your team cannot operate a JVM-based distributed cluster
- you need a fully managed cloud ETL service rather than self-hosted infrastructure
Facets
application · maturity stable
etl streaming message-queue database monitoring rag big-data databases analytics jvm cloud self-hosted cross-platform data-integration cdc change-data-capture elt data-synchronization connectors batch-processing data-ingestion zeta-engine flink spark data-lake schema-evolution exactly-once multimodal-data embeddings llm data-engineering real-time docker kubernetes
10 sources
- readme: https://github.com/apache/seatunnel · fetched 2026-08-28 · fcc4a87000fc
- homepage: https://seatunnel.apache.org/ · fetched 2026-08-29 · 62d89389a531
- site_page: https://seatunnel.apache.org/docs/2.3.13/introduction/about · fetched 2026-08-29 · 41c7294e1981
- site_page: https://seatunnel.apache.org/docs/2.3.12/about · fetched 2026-08-29 · 4425b5991385
- site_page: https://seatunnel.apache.org/docs/2.3.11/about · fetched 2026-08-29 · 2dc65909b72c
- site_page: https://seatunnel.apache.org/docs/2.3.10/about · fetched 2026-08-29 · fc8def483224
- site_page: https://seatunnel.apache.org/docs/2.3.9/about · fetched 2026-08-29 · fc84a091fe2f
- site_page: https://seatunnel.apache.org/docs/introduction/about · fetched 2026-08-29 · 354ba2bbc092
- site_page: https://seatunnel.apache.org/docs/2.3.13/getting-started/locally/quick-start-seatunnel-engine · fetched 2026-08-29 · f47a740c5615
- site_page: https://seatunnel.apache.org/docs/2.3.13/connectors/source · fetched 2026-08-29 · 73d2f4efdd8d
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| apache/seatunnel | main | 82 |
For agents
markdown · JSON · MCP: product_card(name="apache/seatunnel")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem