Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: big-data

477 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Redis
Redis is an open-source, in-memory data structure server written in C that doubles as a cache, NoSQL data store, message broker, and docume…
9576114stable
Pathway
Pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG applications. It provides a unified Py…
9762383active
ClickHouse
ClickHouse is an open-source, column-oriented database management system designed for real-time analytical queries over large datasets. It …
9549470stable
apache/airflow
Apache Airflow is an open-source platform for programmatically authoring, scheduling, and monitoring workflows as directed acyclic graphs (…
9846613stable
apache/spark
Apache Spark is a unified analytics engine for large-scale data processing, providing high-level APIs in Scala, Java, Python, and R over an…
7743882stable
pola-rs/polars
Polars is an extremely fast analytical query engine for DataFrames written in Rust, with multi-threaded vectorized execution, lazy query op…
9539505stable
apache/kafka
Apache Kafka is an open-source distributed event streaming platform for building high-performance data pipelines, streaming analytics, and …
7733632stable
rustfs/rustfs
RustFS is a high-performance, distributed, S3-compatible object storage system written in Rust, positioned as an Apache 2.0-licensed altern…
8031446active
alibaba/canal
Canal is an Alibaba open-source component that parses MySQL binlog to provide incremental data subscription and consumption. It masquerades…
6629725stable
apache/flink
Apache Flink is an open-source distributed stream processing framework for stateful computations over unbounded and bounded data streams, w…
7726294stable
taosdata/TDengine
TDengine is an open-source, cloud-native time-series database (TSDB) written in C, purpose-built for IoT, connected vehicles, industrial Io…
9325090active
apache/rocketmq
Apache RocketMQ is a distributed cloud-native messaging and streaming platform built for low-latency, high-throughput, financially reliable…
9022570stable
airbytehq/airbyte
Airbyte is an open-source data movement platform providing 600+ connectors for replicating data from APIs, databases, and files into wareho…
7921960active
apache/shardingsphere
Apache ShardingSphere is an enterprise distributed database ecosystem that turns heterogeneous databases (MySQL, PostgreSQL, etc.) into a d…
7920788stable
spotify/luigi
Luigi is a Python package for building complex pipelines of long-running batch jobs. It handles dependency resolution, workflow management,…
9118765stable
lightgbm-org/LightGBM
LightGBM is a fast, distributed, high-performance gradient boosting framework based on decision tree algorithms, with APIs for Python, R, C…
8618714stable
alibaba/DataX
DataX is Alibaba's open-source offline data synchronization framework, the open version of Alibaba Cloud DataWorks data integration. It syn…
6417328stable
apache/arrow
Apache Arrow is a language-independent columnar in-memory data format specification plus multi-language libraries (C++, Python/PyArrow, Jav…
9417062stable
prestodb/presto
Presto is a distributed SQL query engine for running interactive queries against large datasets across heterogeneous data sources such as H…
9616724stable
dagster-io/dagster
Dagster is an open-source Python data orchestration platform for developing, producing, and observing data assets, with integrated lineage,…
9516067active
apache/doris
Apache Doris is an open-source MPP-based real-time analytical database that delivers sub-second queries over massive datasets, combining a …
9915817stable
scylladb/scylladb
ScyllaDB is an open-source, high-performance NoSQL wide-column data store built in C++ on the shared-nothing Seastar framework. It is API-c…
7715723stable
apache/hadoop
Apache Hadoop is an open-source framework for reliable, scalable distributed storage (HDFS) and processing (MapReduce, YARN) of large data …
7715640stable
apache/pulsar
Apache Pulsar is a cloud-native, distributed pub-sub messaging and streaming platform originally developed at Yahoo and now a top-level Apa…
9815315stable
elastic/logstash
Logstash is an open-source server-side data processing pipeline that ingests data from multiple sources simultaneously, transforms it, and …
9514925stable
apache/dolphinscheduler
Apache DolphinScheduler is a modern data orchestration platform for building high-performance workflows with low-code drag-and-drop tooling…
9014447stable
juicedata/juicefs
JuiceFS is a high-performance, POSIX-compatible distributed file system that stores file data in object storage (e.g., Amazon S3) and metad…
9814358stable
apache/druid
Apache Druid is a high-performance, distributed, real-time analytics database written in Java for fast OLAP-style slice-and-dice queries on…
8914045stable
Dask
Dask is a flexible parallel and distributed computing library for Python that scales pandas, NumPy, scikit-learn, and other PyData tools to…
9913896stable
opensearch-project/OpenSearch
OpenSearch is an open-source, distributed, RESTful search and analytics suite for full-text search, log analytics, application monitoring, …
9813581stable
trinodb/trino
Trino is a fast, distributed ANSI SQL query engine for big data analytics, formerly known as PrestoSQL. It queries data in place across div…
9313183stable
Debezium
Debezium is an open source distributed platform for change data capture (CDC) that streams row-level database changes as events, most commo…
7713049stable
citusdata/citus
Citus is an open-source PostgreSQL extension (written in C) that transforms Postgres into a distributed database by sharding tables across …
9812730active
redpanda-data/redpanda
Redpanda is a Kafka API-compatible streaming data platform written in C++ on the Seastar framework, with no ZooKeeper or JVM dependency. It…
9412486stable
vesoft-inc/nebula
NebulaGraph is a distributed, open-source graph database written in C++ that handles large graph datasets with millisecond latency and hori…
6012365active
provectus/kafka-ui
UI for Apache Kafka is a free, open-source web UI for monitoring and managing Apache Kafka clusters. It provides a lightweight dashboard to…
2312267active
StarRocks/starrocks
StarRocks is a high-performance distributed OLAP database and query engine built on an MPP architecture with a fully vectorized execution e…
9512043active
manticoresoftware/manticoresearch
Manticore Search is an open-source search database written in C++ that provides fast full-text, vector, and hybrid search with real-time in…
9911956stable
quickwit-oss/quickwit
Quickwit is an open-source, cloud-native search engine written in Rust and powered by Tantivy, designed to run sub-second search and analyt…
9411549active
AutoMQ/automq
AutoMQ is a cloud-native, 100% Apache Kafka-compatible streaming platform that replaces broker-local disks with S3-compatible object storag…
9510574active
modin-project/modin
Modin is a drop-in replacement for pandas that scales DataFrame operations across all CPU cores using Ray, Dask, or Unidist as execution en…
6710390active
oceanbase/oceanbase
OceanBase is a distributed relational database engine developed by Ant Group, offering MySQL compatibility, Paxos-based high availability, …
9810255stable
apache/cassandra
Apache Cassandra is an open-source NoSQL distributed wide-column database written in Java, designed for linear scalability and fault tolera…
7710082stable
NVIDIA/cudf
cuDF is a GPU-accelerated DataFrame library for tabular data processing, part of NVIDIA's RAPIDS suite. It provides a pandas-compatible Pyt…
959734stable
apache/seatunnel
Apache SeaTunnel is a distributed, high-performance data integration platform for synchronizing massive amounts of data across hundreds of …
829587stable
databendlabs/databend
Databend is an open-source, cloud-native data warehouse built in Rust that runs entirely on object storage (S3, Azure, GCS). It unifies BI …
939423active
apache/datafusion
Apache DataFusion is an extensible query engine written in Rust that uses Apache Arrow as its in-memory columnar format. It provides SQL an…
779202active
apache/iceberg
Apache Iceberg is a high-performance open table format for huge analytic datasets, bringing SQL table reliability to big data lakes. This r…
909177stable
Delta Lake
Delta Lake is an open-source storage framework and table format that brings ACID transactions, scalable metadata handling, schema enforceme…
958960stable
redpanda-data/connect
Redpanda Connect (formerly Benthos) is a declarative stream processor that moves data between hundreds of sources and sinks with transforma…
958736active
apache/beam
Apache Beam is an open-source unified programming model and SDK set (Java, Python, Go, SQL, TypeScript) for defining batch and streaming da…
938650stable
pentaho/pentaho-kettle
Pentaho Data Integration (Kettle) is an open-source ETL application for designing and running data extraction, transformation, and loading …
678382active
h2oai/h2o-3
H2O-3 is an open-source, distributed, in-memory machine learning platform implementing algorithms such as GLM, GBM/XGBoost, Random Forest, …
777494stable
arkime/arkime
Arkime is an open-source, large-scale network analysis, full packet capture, and session indexing system that stores traffic in standard PC…
997459active
Alluxio/alluxio
Alluxio is an open-source distributed caching and data orchestration platform that sits between compute frameworks (Spark, Presto, Trino, P…
317231stable
feast-dev/feast
Feast is an open-source feature store for machine learning that manages offline stores for historical training data and low-latency online …
997230stable
didi/KnowStreaming
Know Streaming is a cloud-native Kafka management and control platform built from years of operational experience at Didi. It provides zero…
877177active
vespa-engine/vespa
Vespa is an open-source, distributed AI search platform and serving engine that combines full-text search, vector/tensor search, and machin…
957069stable
snowplow/snowplow
Snowplow is a Customer Data Infrastructure platform that collects, validates, enriches, and streams event-level behavioral data from web, m…
637029stable
apache/zeppelin
Apache Zeppelin is a web-based notebook for interactive, data-driven analytics and collaborative documents. It supports SQL, Scala, Python,…
776655stable
Hazelcast
Hazelcast is a unified real-time data platform combining distributed stream processing with a fast, in-memory data store. It lets applicati…
826604stable
apache/flink-cdc
Flink CDC is a distributed streaming data integration tool built on Apache Flink that captures change data from databases like MySQL and Po…
836468active
pachyderm/pachyderm
Pachyderm is a data-centric pipeline platform that automates data transformations with built-in data versioning and lineage tracking. It ru…
366308active
apache/hudi
Apache Hudi is an open data lakehouse platform built on a high-performance open table format that brings database functionality like transa…
916219stable
apache/nifi
Apache NiFi is an easy-to-use, powerful, and reliable system to process and distribute data, built on the JVM. It provides a browser-based …
946208stable
apache/pinot
Apache Pinot is an open-source distributed OLAP datastore purpose-built for low-latency, high-throughput real-time analytics. It ingests da…
846128stable
OpenAtomFoundation/pikiwidb
PikiwiDB (Pika) is a Redis-compatible, persistent key-value database built on RocksDB, developed by Qihoo's infrastructure team. It speaks …
916123active
WeiYe-Jing/datax-web
DataX-Web is a distributed data synchronization tool built on top of Alibaba's DataX, providing a web UI to visually configure and manage d…
236018stable
apache/hive
Apache Hive is a distributed, fault-tolerant data warehouse system that enables reading, writing, and managing petabytes of data in distrib…
776015active
JanusGraph/janusgraph
JanusGraph is an open-source, distributed graph database optimized for storing and querying graphs with billions of vertices and edges acro…
785829active
dlt-hub/dlt
dlt (data load tool) is an open-source Python library for building ELT data pipelines that extract data from REST APIs, SQL databases, clou…
985778stable
nuclio/nuclio
Nuclio is a high-performance open-source serverless (FaaS) platform for real-time event and data processing, deployable standalone via Dock…
955750stable
Eventual-Inc/Daft
Daft is a high-performance distributed data engine with a Python dataframe API, implemented in Rust, designed for AI and multimodal workloa…
955730active
cubefs/cubefs
CubeFS is a CNCF-graduated, cloud-native distributed file and object storage system written in Go. It supports S3, HDFS, and POSIX access p…
905636stable
apache/hbase
Apache HBase is an open-source, distributed, versioned, column-oriented NoSQL database modeled after Google Bigtable, running on top of Apa…
945552stable
treeverse/lakeFS
lakeFS is an open-source data version control service that turns object storage (S3, Azure Blob, GCS) into a Git-like repository with branc…
945496stable
fluvio-community/fluvio
Fluvio is a distributed data streaming engine written in Rust that combines a Kafka-like event streaming platform with the Stateful DataFlo…
795247active
microsoft/SynapseML
SynapseML (formerly MMLSpark) is an open-source machine learning library built on Apache Spark that provides simple, composable, distribute…
885240active
apache/calcite
Apache Calcite is a dynamic data management framework for the JVM that supplies the building blocks of a database—an industry-standard SQL …
775175stable
apache/ignite
Apache Ignite is a distributed, memory-first database with in-memory speed, ACID transactions, and SQL support across a cluster. It can ser…
775079stable
jitsucom/jitsu
Jitsu is an open-source, self-hostable event data platform and Segment alternative that collects event data from websites, apps, and server…
975043active
ArroyoSystems/arroyo
Arroyo is a distributed stream processing engine written in Rust that lets users run stateful computations on high-volume real-time data st…
785016active
microsoft/SPTAG
SPTAG is a C++ library from Microsoft for large-scale approximate nearest neighbor (ANN) vector search, offering kd-tree/balanced k-means t…
765012active
deepseek-ai/smallpond
Smallpond is a lightweight distributed data processing framework built on DuckDB and DeepSeek's 3FS shared file system. It lets users proce…
245000active
apache/age
Apache AGE is a PostgreSQL extension that adds graph database capabilities on top of existing relational databases, supporting openCypher q…
934784active
amundsen-io/amundsen
Amundsen is an open-source data discovery and metadata engine that indexes data resources such as tables, dashboards, and streams, and powe…
674782active
ydb-platform/ydb
YDB is an open-source distributed SQL DBMS written in C++ that combines horizontal scalability, high availability, and strict consistency w…
984770stable
apache/rocketmq-externals
The Apache RocketMQ externals repository is the community home for incubating ecosystem projects around the RocketMQ message broker, such a…
714601active
rudderlabs/rudder-server
RudderStack's open-source event streaming server, a privacy- and security-focused Segment alternative written in Go. It collects customer e…
954477active
rom1504/img2dataset
A Python tool that downloads large sets of image URLs and packages them into machine learning datasets, with resizing and caption support. …
564443active
crate/crate
CrateDB is a distributed, horizontally scalable SQL database built on Lucene, designed for real-time analytics on massive datasets. It supp…
954427active
apache/streampark
Apache StreamPark is a streaming application development framework and one-stop cloud-native real-time computing platform for Apache Flink …
734328stable
zendesk/maxwell
Maxwell's Daemon is a change data capture (CDC) application that reads MySQL binlogs and emits row-level changes as JSON to Kafka, Kinesis,…
964258active
facebookincubator/velox
Velox is a composable, extensible C++ execution engine library for data management systems, created by Meta. It provides reusable vectorize…
774201active
pydata/xarray
Xarray is a Python library that adds labeled dimensions, coordinates, and attributes on top of NumPy-like N-dimensional arrays and datasets…
974190stable
DTStack/chunjun
ChunJun (formerly FlinkX) is a distributed, batch-and-stream data integration framework built on Apache Flink. It synchronizes and computes…
544099active
Roaring Bitmaps
Roaring Bitmaps is a Java library providing compressed bitsets that outperform conventional compressed bitmap formats like WAH, EWAH, and C…
973918stable
Rdatatable/data.table
data.table is an R package providing a high-performance, memory-efficient replacement for base R's data.frame with a concise indexing synta…
773913stable
Netflix/maestro
Netflix Maestro is a general-purpose workflow orchestrator providing workflow-as-a-service for data, ML, and software pipelines. It schedul…
693829active
apache/kylin
Apache Kylin is an open-source distributed OLAP engine for big data that delivers sub-second query latency on trillions of records. It prov…
643773stable

page 1 / 5 next →