domain: big-data
477 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Redis Redis is an open-source, in-memory data structure server written in C that doubles as a cache, NoSQL data store, message broker, and docume… | 95 | 76114 | stable |
| Pathway Pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG applications. It provides a unified Py… | 97 | 62383 | active |
| ClickHouse ClickHouse is an open-source, column-oriented database management system designed for real-time analytical queries over large datasets. It … | 95 | 49470 | stable |
| apache/airflow Apache Airflow is an open-source platform for programmatically authoring, scheduling, and monitoring workflows as directed acyclic graphs (… | 98 | 46613 | stable |
| apache/spark Apache Spark is a unified analytics engine for large-scale data processing, providing high-level APIs in Scala, Java, Python, and R over an… | 77 | 43882 | stable |
| pola-rs/polars Polars is an extremely fast analytical query engine for DataFrames written in Rust, with multi-threaded vectorized execution, lazy query op… | 95 | 39505 | stable |
| apache/kafka Apache Kafka is an open-source distributed event streaming platform for building high-performance data pipelines, streaming analytics, and … | 77 | 33632 | stable |
| rustfs/rustfs RustFS is a high-performance, distributed, S3-compatible object storage system written in Rust, positioned as an Apache 2.0-licensed altern… | 80 | 31446 | active |
| alibaba/canal Canal is an Alibaba open-source component that parses MySQL binlog to provide incremental data subscription and consumption. It masquerades… | 66 | 29725 | stable |
| apache/flink Apache Flink is an open-source distributed stream processing framework for stateful computations over unbounded and bounded data streams, w… | 77 | 26294 | stable |
| taosdata/TDengine TDengine is an open-source, cloud-native time-series database (TSDB) written in C, purpose-built for IoT, connected vehicles, industrial Io… | 93 | 25090 | active |
| apache/rocketmq Apache RocketMQ is a distributed cloud-native messaging and streaming platform built for low-latency, high-throughput, financially reliable… | 90 | 22570 | stable |
| airbytehq/airbyte Airbyte is an open-source data movement platform providing 600+ connectors for replicating data from APIs, databases, and files into wareho… | 79 | 21960 | active |
| apache/shardingsphere Apache ShardingSphere is an enterprise distributed database ecosystem that turns heterogeneous databases (MySQL, PostgreSQL, etc.) into a d… | 79 | 20788 | stable |
| spotify/luigi Luigi is a Python package for building complex pipelines of long-running batch jobs. It handles dependency resolution, workflow management,… | 91 | 18765 | stable |
| lightgbm-org/LightGBM LightGBM is a fast, distributed, high-performance gradient boosting framework based on decision tree algorithms, with APIs for Python, R, C… | 86 | 18714 | stable |
| alibaba/DataX DataX is Alibaba's open-source offline data synchronization framework, the open version of Alibaba Cloud DataWorks data integration. It syn… | 64 | 17328 | stable |
| apache/arrow Apache Arrow is a language-independent columnar in-memory data format specification plus multi-language libraries (C++, Python/PyArrow, Jav… | 94 | 17062 | stable |
| prestodb/presto Presto is a distributed SQL query engine for running interactive queries against large datasets across heterogeneous data sources such as H… | 96 | 16724 | stable |
| dagster-io/dagster Dagster is an open-source Python data orchestration platform for developing, producing, and observing data assets, with integrated lineage,… | 95 | 16067 | active |
| apache/doris Apache Doris is an open-source MPP-based real-time analytical database that delivers sub-second queries over massive datasets, combining a … | 99 | 15817 | stable |
| scylladb/scylladb ScyllaDB is an open-source, high-performance NoSQL wide-column data store built in C++ on the shared-nothing Seastar framework. It is API-c… | 77 | 15723 | stable |
| apache/hadoop Apache Hadoop is an open-source framework for reliable, scalable distributed storage (HDFS) and processing (MapReduce, YARN) of large data … | 77 | 15640 | stable |
| apache/pulsar Apache Pulsar is a cloud-native, distributed pub-sub messaging and streaming platform originally developed at Yahoo and now a top-level Apa… | 98 | 15315 | stable |
| elastic/logstash Logstash is an open-source server-side data processing pipeline that ingests data from multiple sources simultaneously, transforms it, and … | 95 | 14925 | stable |
| apache/dolphinscheduler Apache DolphinScheduler is a modern data orchestration platform for building high-performance workflows with low-code drag-and-drop tooling… | 90 | 14447 | stable |
| juicedata/juicefs JuiceFS is a high-performance, POSIX-compatible distributed file system that stores file data in object storage (e.g., Amazon S3) and metad… | 98 | 14358 | stable |
| apache/druid Apache Druid is a high-performance, distributed, real-time analytics database written in Java for fast OLAP-style slice-and-dice queries on… | 89 | 14045 | stable |
| Dask Dask is a flexible parallel and distributed computing library for Python that scales pandas, NumPy, scikit-learn, and other PyData tools to… | 99 | 13896 | stable |
| opensearch-project/OpenSearch OpenSearch is an open-source, distributed, RESTful search and analytics suite for full-text search, log analytics, application monitoring, … | 98 | 13581 | stable |
| trinodb/trino Trino is a fast, distributed ANSI SQL query engine for big data analytics, formerly known as PrestoSQL. It queries data in place across div… | 93 | 13183 | stable |
| Debezium Debezium is an open source distributed platform for change data capture (CDC) that streams row-level database changes as events, most commo… | 77 | 13049 | stable |
| citusdata/citus Citus is an open-source PostgreSQL extension (written in C) that transforms Postgres into a distributed database by sharding tables across … | 98 | 12730 | active |
| redpanda-data/redpanda Redpanda is a Kafka API-compatible streaming data platform written in C++ on the Seastar framework, with no ZooKeeper or JVM dependency. It… | 94 | 12486 | stable |
| vesoft-inc/nebula NebulaGraph is a distributed, open-source graph database written in C++ that handles large graph datasets with millisecond latency and hori… | 60 | 12365 | active |
| provectus/kafka-ui UI for Apache Kafka is a free, open-source web UI for monitoring and managing Apache Kafka clusters. It provides a lightweight dashboard to… | 23 | 12267 | active |
| StarRocks/starrocks StarRocks is a high-performance distributed OLAP database and query engine built on an MPP architecture with a fully vectorized execution e… | 95 | 12043 | active |
| manticoresoftware/manticoresearch Manticore Search is an open-source search database written in C++ that provides fast full-text, vector, and hybrid search with real-time in… | 99 | 11956 | stable |
| quickwit-oss/quickwit Quickwit is an open-source, cloud-native search engine written in Rust and powered by Tantivy, designed to run sub-second search and analyt… | 94 | 11549 | active |
| AutoMQ/automq AutoMQ is a cloud-native, 100% Apache Kafka-compatible streaming platform that replaces broker-local disks with S3-compatible object storag… | 95 | 10574 | active |
| modin-project/modin Modin is a drop-in replacement for pandas that scales DataFrame operations across all CPU cores using Ray, Dask, or Unidist as execution en… | 67 | 10390 | active |
| oceanbase/oceanbase OceanBase is a distributed relational database engine developed by Ant Group, offering MySQL compatibility, Paxos-based high availability, … | 98 | 10255 | stable |
| apache/cassandra Apache Cassandra is an open-source NoSQL distributed wide-column database written in Java, designed for linear scalability and fault tolera… | 77 | 10082 | stable |
| NVIDIA/cudf cuDF is a GPU-accelerated DataFrame library for tabular data processing, part of NVIDIA's RAPIDS suite. It provides a pandas-compatible Pyt… | 95 | 9734 | stable |
| apache/seatunnel Apache SeaTunnel is a distributed, high-performance data integration platform for synchronizing massive amounts of data across hundreds of … | 82 | 9587 | stable |
| databendlabs/databend Databend is an open-source, cloud-native data warehouse built in Rust that runs entirely on object storage (S3, Azure, GCS). It unifies BI … | 93 | 9423 | active |
| apache/datafusion Apache DataFusion is an extensible query engine written in Rust that uses Apache Arrow as its in-memory columnar format. It provides SQL an… | 77 | 9202 | active |
| apache/iceberg Apache Iceberg is a high-performance open table format for huge analytic datasets, bringing SQL table reliability to big data lakes. This r… | 90 | 9177 | stable |
| Delta Lake Delta Lake is an open-source storage framework and table format that brings ACID transactions, scalable metadata handling, schema enforceme… | 95 | 8960 | stable |
| redpanda-data/connect Redpanda Connect (formerly Benthos) is a declarative stream processor that moves data between hundreds of sources and sinks with transforma… | 95 | 8736 | active |
| apache/beam Apache Beam is an open-source unified programming model and SDK set (Java, Python, Go, SQL, TypeScript) for defining batch and streaming da… | 93 | 8650 | stable |
| pentaho/pentaho-kettle Pentaho Data Integration (Kettle) is an open-source ETL application for designing and running data extraction, transformation, and loading … | 67 | 8382 | active |
| h2oai/h2o-3 H2O-3 is an open-source, distributed, in-memory machine learning platform implementing algorithms such as GLM, GBM/XGBoost, Random Forest, … | 77 | 7494 | stable |
| arkime/arkime Arkime is an open-source, large-scale network analysis, full packet capture, and session indexing system that stores traffic in standard PC… | 99 | 7459 | active |
| Alluxio/alluxio Alluxio is an open-source distributed caching and data orchestration platform that sits between compute frameworks (Spark, Presto, Trino, P… | 31 | 7231 | stable |
| feast-dev/feast Feast is an open-source feature store for machine learning that manages offline stores for historical training data and low-latency online … | 99 | 7230 | stable |
| didi/KnowStreaming Know Streaming is a cloud-native Kafka management and control platform built from years of operational experience at Didi. It provides zero… | 87 | 7177 | active |
| vespa-engine/vespa Vespa is an open-source, distributed AI search platform and serving engine that combines full-text search, vector/tensor search, and machin… | 95 | 7069 | stable |
| snowplow/snowplow Snowplow is a Customer Data Infrastructure platform that collects, validates, enriches, and streams event-level behavioral data from web, m… | 63 | 7029 | stable |
| apache/zeppelin Apache Zeppelin is a web-based notebook for interactive, data-driven analytics and collaborative documents. It supports SQL, Scala, Python,… | 77 | 6655 | stable |
| Hazelcast Hazelcast is a unified real-time data platform combining distributed stream processing with a fast, in-memory data store. It lets applicati… | 82 | 6604 | stable |
| apache/flink-cdc Flink CDC is a distributed streaming data integration tool built on Apache Flink that captures change data from databases like MySQL and Po… | 83 | 6468 | active |
| pachyderm/pachyderm Pachyderm is a data-centric pipeline platform that automates data transformations with built-in data versioning and lineage tracking. It ru… | 36 | 6308 | active |
| apache/hudi Apache Hudi is an open data lakehouse platform built on a high-performance open table format that brings database functionality like transa… | 91 | 6219 | stable |
| apache/nifi Apache NiFi is an easy-to-use, powerful, and reliable system to process and distribute data, built on the JVM. It provides a browser-based … | 94 | 6208 | stable |
| apache/pinot Apache Pinot is an open-source distributed OLAP datastore purpose-built for low-latency, high-throughput real-time analytics. It ingests da… | 84 | 6128 | stable |
| OpenAtomFoundation/pikiwidb PikiwiDB (Pika) is a Redis-compatible, persistent key-value database built on RocksDB, developed by Qihoo's infrastructure team. It speaks … | 91 | 6123 | active |
| WeiYe-Jing/datax-web DataX-Web is a distributed data synchronization tool built on top of Alibaba's DataX, providing a web UI to visually configure and manage d… | 23 | 6018 | stable |
| apache/hive Apache Hive is a distributed, fault-tolerant data warehouse system that enables reading, writing, and managing petabytes of data in distrib… | 77 | 6015 | active |
| JanusGraph/janusgraph JanusGraph is an open-source, distributed graph database optimized for storing and querying graphs with billions of vertices and edges acro… | 78 | 5829 | active |
| dlt-hub/dlt dlt (data load tool) is an open-source Python library for building ELT data pipelines that extract data from REST APIs, SQL databases, clou… | 98 | 5778 | stable |
| nuclio/nuclio Nuclio is a high-performance open-source serverless (FaaS) platform for real-time event and data processing, deployable standalone via Dock… | 95 | 5750 | stable |
| Eventual-Inc/Daft Daft is a high-performance distributed data engine with a Python dataframe API, implemented in Rust, designed for AI and multimodal workloa… | 95 | 5730 | active |
| cubefs/cubefs CubeFS is a CNCF-graduated, cloud-native distributed file and object storage system written in Go. It supports S3, HDFS, and POSIX access p… | 90 | 5636 | stable |
| apache/hbase Apache HBase is an open-source, distributed, versioned, column-oriented NoSQL database modeled after Google Bigtable, running on top of Apa… | 94 | 5552 | stable |
| treeverse/lakeFS lakeFS is an open-source data version control service that turns object storage (S3, Azure Blob, GCS) into a Git-like repository with branc… | 94 | 5496 | stable |
| fluvio-community/fluvio Fluvio is a distributed data streaming engine written in Rust that combines a Kafka-like event streaming platform with the Stateful DataFlo… | 79 | 5247 | active |
| microsoft/SynapseML SynapseML (formerly MMLSpark) is an open-source machine learning library built on Apache Spark that provides simple, composable, distribute… | 88 | 5240 | active |
| apache/calcite Apache Calcite is a dynamic data management framework for the JVM that supplies the building blocks of a database—an industry-standard SQL … | 77 | 5175 | stable |
| apache/ignite Apache Ignite is a distributed, memory-first database with in-memory speed, ACID transactions, and SQL support across a cluster. It can ser… | 77 | 5079 | stable |
| jitsucom/jitsu Jitsu is an open-source, self-hostable event data platform and Segment alternative that collects event data from websites, apps, and server… | 97 | 5043 | active |
| ArroyoSystems/arroyo Arroyo is a distributed stream processing engine written in Rust that lets users run stateful computations on high-volume real-time data st… | 78 | 5016 | active |
| microsoft/SPTAG SPTAG is a C++ library from Microsoft for large-scale approximate nearest neighbor (ANN) vector search, offering kd-tree/balanced k-means t… | 76 | 5012 | active |
| deepseek-ai/smallpond Smallpond is a lightweight distributed data processing framework built on DuckDB and DeepSeek's 3FS shared file system. It lets users proce… | 24 | 5000 | active |
| apache/age Apache AGE is a PostgreSQL extension that adds graph database capabilities on top of existing relational databases, supporting openCypher q… | 93 | 4784 | active |
| amundsen-io/amundsen Amundsen is an open-source data discovery and metadata engine that indexes data resources such as tables, dashboards, and streams, and powe… | 67 | 4782 | active |
| ydb-platform/ydb YDB is an open-source distributed SQL DBMS written in C++ that combines horizontal scalability, high availability, and strict consistency w… | 98 | 4770 | stable |
| apache/rocketmq-externals The Apache RocketMQ externals repository is the community home for incubating ecosystem projects around the RocketMQ message broker, such a… | 71 | 4601 | active |
| rudderlabs/rudder-server RudderStack's open-source event streaming server, a privacy- and security-focused Segment alternative written in Go. It collects customer e… | 95 | 4477 | active |
| rom1504/img2dataset A Python tool that downloads large sets of image URLs and packages them into machine learning datasets, with resizing and caption support. … | 56 | 4443 | active |
| crate/crate CrateDB is a distributed, horizontally scalable SQL database built on Lucene, designed for real-time analytics on massive datasets. It supp… | 95 | 4427 | active |
| apache/streampark Apache StreamPark is a streaming application development framework and one-stop cloud-native real-time computing platform for Apache Flink … | 73 | 4328 | stable |
| zendesk/maxwell Maxwell's Daemon is a change data capture (CDC) application that reads MySQL binlogs and emits row-level changes as JSON to Kafka, Kinesis,… | 96 | 4258 | active |
| facebookincubator/velox Velox is a composable, extensible C++ execution engine library for data management systems, created by Meta. It provides reusable vectorize… | 77 | 4201 | active |
| pydata/xarray Xarray is a Python library that adds labeled dimensions, coordinates, and attributes on top of NumPy-like N-dimensional arrays and datasets… | 97 | 4190 | stable |
| DTStack/chunjun ChunJun (formerly FlinkX) is a distributed, batch-and-stream data integration framework built on Apache Flink. It synchronizes and computes… | 54 | 4099 | active |
| Roaring Bitmaps Roaring Bitmaps is a Java library providing compressed bitsets that outperform conventional compressed bitmap formats like WAH, EWAH, and C… | 97 | 3918 | stable |
| Rdatatable/data.table data.table is an R package providing a high-performance, memory-efficient replacement for base R's data.frame with a concise indexing synta… | 77 | 3913 | stable |
| Netflix/maestro Netflix Maestro is a general-purpose workflow orchestrator providing workflow-as-a-service for data, ML, and software pipelines. It schedul… | 69 | 3829 | active |
| apache/kylin Apache Kylin is an open-source distributed OLAP engine for big data that delivers sub-second query latency on trillions of records. It prov… | 64 | 3773 | stable |
page 1 / 5 next →