domain: big-data
477 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| DataLinkDC/dinky Dinky is an open-source, one-stop real-time computing platform built around Apache Flink, providing an IDE for FlinkSQL development, online… | 80 | 3752 | active |
| awslabs/deequ Deequ is a Scala library built on Apache Spark for defining 'unit tests for data' that measure data quality in large datasets. It computes … | 93 | 3643 | active |
| TimelyDataflow/timely-dataflow A modular Rust implementation of the timely dataflow computational model from the Naiad paper, providing a low-latency cyclic data-parallel… | 92 | 3642 | active |
| apache/arrow-rs The official Rust implementation of Apache Arrow, providing the in-memory columnar format, and of Apache Parquet for columnar file reading … | 98 | 3592 | active |
| Netflix/atlas Atlas is Netflix's in-memory dimensional time series database, serving as a backend for storing and querying operational metrics at scale. … | 90 | 3564 | active |
| chrislusf/gleam Gleam is a fast, efficient distributed map/reduce execution system written in pure Go, defining computations as DAG flows that can run stan… | 75 | 3562 | active |
| holoviz/datashader Datashader is a Python data rasterization pipeline that renders very large datasets into fixed-size images by projecting, aggregating, and … | 89 | 3558 | active |
| alibaba/GraphScope GraphScope is a unified distributed graph computing platform from Alibaba that combines graph analytics (GRAPE), interactive graph queries … | 73 | 3556 | active |
| NLPIR-team/NLPIR NLPIR is a Chinese natural language processing platform (formerly ICTCLAS) providing an SDK with around twenty text analysis functions incl… | 75 | 3474 | active |
| apache/datafusion-sqlparser-rs An extensible SQL lexer and parser library for Rust that produces an AST conforming to the ANSI/ISO SQL standard and various vendor dialect… | 77 | 3432 | active |
| apache/linkis Apache Linkis is a computation middleware layer that sits between upper-layer applications and underlying big data engines such as Spark, H… | 71 | 3406 | stable |
| apache/paimon Apache Paimon is a lakehouse table format that combines lake format with LSM tree structure to support real-time streaming updates. It enab… | 87 | 3384 | active |
| lakehq/sail Sail is an open-source, Rust-native multimodal compute engine that serves as a drop-in replacement for Apache Spark, unifying batch process… | 93 | 3333 | active |
| apache/avro Apache Avro is a data serialization system providing rich data structures, a compact binary format, container files, and RPC support, with … | 88 | 3301 | stable |
| root-project/root ROOT is a C++ data analysis framework developed at CERN for storing, processing, and visualizing large scientific datasets, with a columnar… | 96 | 3288 | stable |
| apache/nutch Apache Nutch is a highly extensible and scalable open-source web crawler built on Apache Hadoop data structures. It supports batch crawling… | 77 | 3277 | stable |
| WeBankFinTech/DataSphereStudio DataSphere Studio is a one-stop data application development and management portal from WeBank, built on the Linkis computing middleware. I… | 50 | 3262 | active |
| PeerDB-io/peerdb PeerDB is an open-source ETL/CDC tool that streams data from Postgres to data warehouses, queues, and storage engines, with an actively mai… | 93 | 3252 | active |
| lakesoul-io/LakeSoul LakeSoul is a cloud-native, real-time lakehouse framework with a Rust-native core providing ACID table format, concurrent upserts, incremen… | 82 | 3247 | active |
| duckdb/pg_duckdb pg_duckdb is an official PostgreSQL extension that embeds DuckDB's columnar-vectorized analytics engine inside Postgres. It lets users run … | 72 | 3210 | active |
| apache/gravitino Apache Gravitino is a high-performance, geo-distributed, federated metadata lake that provides unified metadata management across diverse d… | 89 | 3188 | active |
| vortex-data/vortex Vortex is an extensible, high-performance columnar file format and toolkit for compressed data, positioned as a faster alternative to Apach… | 88 | 3158 | active |
| apache/hugegraph Apache HugeGraph is a fast, highly scalable graph database supporting tens of billions of vertices and edges with OLTP real-time queries vi… | 73 | 3157 | active |
| Tencent/Tendis Tendis is a high-performance distributed key-value storage system fully compatible with the Redis protocol, built on RocksDB for disk-based… | 81 | 3155 | active |
| blockchain-etl/ethereum-etl A Python CLI tool that extracts Ethereum blockchain data (blocks, transactions, token transfers, logs, traces, contracts) and exports it to… | 52 | 3127 | active |
| alldatacenter/alldata AllData is a definable data platform (数据中台) that integrates open-source components like DolphinScheduler, DataX, SeaTunnel, OpenMetadata, a… | 83 | 3082 | active |
| apache/parquet-java Apache Parquet Java is the Java implementation of the Apache Parquet column-oriented data file format. It provides reading and writing of e… | 90 | 3076 | active |
| heavyai/heavydb HeavyDB (formerly MapD/OmniSciDB) is an open-source SQL-based, relational, columnar database engine that uses CPUs and NVIDIA GPUs to query… | 72 | 3057 | active |
| hydradatabase/columnar Hydra Columnar is a Postgres extension providing column-oriented table storage via the table access method API, enabling fast analytical qu… | 26 | 3039 | stable |
| TimelyDataflow/differential-dataflow A Rust library implementing differential dataflow on top of timely dataflow, enabling data-parallel programs that efficiently process large… | 93 | 2998 | active |
| ekzhu/datasketch datasketch is a Python library of probabilistic data structures (MinHash, Weighted MinHash, HyperLogLog, HyperLogLog++) and indexes (MinHas… | 91 | 2959 | stable |
| duckdb/ducklake DuckLake is an integrated data lake and lakehouse catalog format that stores data as Parquet files and manages metadata in an ACID-complian… | 65 | 2939 | active |
| yahoo/CMAK CMAK (Cluster Manager for Apache Kafka, formerly Kafka Manager) is a web-based tool for managing and inspecting Apache Kafka clusters. It s… | 23 | 11925 | maintenance |
| numaproj/numaflow Numaflow is a Kubernetes-native, serverless platform for running massively parallel data and streaming processing jobs. It lets developers … | 98 | 2825 | active |
| kafbat/kafka-ui Kafbat UI is a free, open-source web UI for monitoring and managing Apache Kafka clusters, supporting multi-cluster management, topic and m… | 78 | 2635 | active |
| spotify/scio Scio is a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. It provides a unified batch and strea… | 96 | 2628 | active |
| OpenLineage/OpenLineage OpenLineage is an open standard and framework for collecting data lineage metadata, defining a generic model of jobs, runs, and datasets wi… | 97 | 2626 | active |
| meltano/meltano Meltano is an open-source, declarative, code-first data integration engine and CLI for building and running ELT pipelines. It orchestrates … | 97 | 2610 | stable |
| pingcap/ossinsight OSSInsight is a web analytics platform that analyzes over 10 billion GitHub events to provide rankings, trends, and comparisons of open sou… | 67 | 2498 | active |
| griddb/griddb GridDB is an open-source, high-performance distributed database optimized for time-series IoT and big data workloads, offering both NoSQL a… | 66 | 2475 | active |
| confluentinc/schema-registry Confluent Schema Registry is a Java-based serving layer that stores and retrieves Avro, JSON Schema, and Protobuf schemas via a RESTful API… | 77 | 2462 | active |
| Mycat MyCAT is an open-source database middleware written in Java that acts as a distributed MySQL cluster proxy, providing sharding, read/write … | 60 | 9514 | maintenance |
| elementary-data/elementary Elementary OSS is an open-source CLI for dbt-native data observability that reads warehouse metadata, dbt artifacts, and test results to ge… | 97 | 2398 | active |
| apache/sedona Apache Sedona is a cluster computing framework for processing large-scale geospatial data across Spark, Flink, and Snowflake. It provides S… | 90 | 2390 | active |
| apache/geode Apache Geode is a distributed in-memory data management platform (originally GemFire) providing real-time, consistent, low-latency access t… | 90 | 2378 | stable |
| quarylabs/quary Quary is an open-source business intelligence tool for engineers that connects to databases, lets users write SQL to transform and document… | 81 | 2376 | active |
| syslog-ng/syslog-ng syslog-ng is an enhanced syslog daemon that collects, parses, filters, and routes log messages from a wide range of sources including syslo… | 92 | 2371 | stable |
| moj-analytical-services/splink Splink is a Python library for fast, scalable probabilistic record linkage and entity resolution, based on the Fellegi-Sunter model with un… | 90 | 2364 | active |
| apache/kyuubi Apache Kyuubi is a distributed, multi-tenant Thrift JDBC/ODBC gateway that provides serverless SQL access to data warehouses and lakehouses… | 94 | 2361 | stable |
| apache/ambari Apache Ambari is an open-source tool for provisioning, managing, and monitoring Apache Hadoop clusters. It provides a browser-based managem… | 79 | 2311 | active |
| twitter/algebird Algebird is a Scala library providing abstract algebra typeclasses (Monoids, Groups, Rings) for building composable aggregation systems. It… | 47 | 2295 | stable |
| pinterest/querybook Querybook is an open-source big data IDE from Pinterest that combines a collaborative notebook interface (DataDocs) with table metadata fro… | 67 | 2280 | active |
| apache/gobblin Apache Gobblin is a distributed data integration framework for ingesting, replicating, organizing, and managing lifecycle of data across st… | 66 | 2270 | stable |
| MarquezProject/marquez Marquez is an open-source metadata service for collecting, aggregating, and visualizing a data ecosystem's metadata. It serves as the refer… | 67 | 2268 | active |
| timeplus-io/proton Timeplus Proton is a unified streaming SQL engine shipped as a single dependency-free C++ binary, built on the ClickHouse engine. It ingest… | 93 | 2247 | active |
| nathanmarz/storm Apache Storm (originally nathanmarz/storm) is a distributed, fault-tolerant realtime computation system for stream processing, continuous c… | 32 | 8765 | maintenance |
| ytsaurus/ytsaurus YTsaurus is an open-source, fault-tolerant big data platform combining distributed storage, MapReduce processing, a SQL query engine, and a… | 94 | 2200 | active |
| vaexio/vaex Vaex is a high-performance Python DataFrame library for lazy, out-of-core processing of large tabular datasets, using memory mapping and ze… | 57 | 8507 | maintenance |
| reugn/go-streams A lightweight stream processing library for Go offering a concise DSL to build declarative data pipelines from composable sources, flows, a… | 58 | 2170 | active |
| fugue-project/fugue Fugue is a Python library providing a unified interface for distributed computing, letting users run Python, Pandas, Polars, and SQL code o… | 82 | 2169 | active |
| tdunning/t-digest A Java library implementing the t-digest data structure for accurate online accumulation of rank-based statistics such as quantiles and tri… | 35 | 2166 | stable |
| apache/atlas Apache Atlas is an open-source metadata management and data governance framework providing a shared metadata store, data lineage, and class… | 77 | 2136 | active |
| soedinglab/MMseqs2 MMseqs2 is an open-source C++ suite for ultra-fast, sensitive search and clustering of huge protein and nucleotide sequence sets, running o… | 70 | 2131 | active |
| apache/datafusion-ballista Apache DataFusion Ballista is a distributed query execution engine built on Apache DataFusion that parallelizes SQL and DataFrame workloads… | 98 | 2115 | active |
| apache/fluss Apache Fluss is a lakehouse-native streaming storage built for real-time analytics and AI, serving as the real-time data layer for Lakehous… | 78 | 2113 | active |
| warpstreamlabs/bento Bento is a high-performance, resilient stream processor written in Go that connects a wide range of sources and sinks (Kafka, Pub/Sub, Redi… | 91 | 2112 | active |
| TileDB-Inc/TileDB TileDB is an embeddable C++ storage engine for dense and sparse multi-dimensional arrays, with cloud object storage support (S3, GCS, Azure… | 86 | 2073 | stable |
| feldera/feldera Feldera is an incremental computation engine written in Rust that evaluates arbitrary SQL programs incrementally using DBSP theory, maintai… | 92 | 2064 | active |
| zarr-developers/zarr-python Zarr is a Python library implementing the Zarr storage format for chunked, compressed, N-dimensional arrays with NumPy-compatible dtypes. I… | 98 | 2044 | stable |
| apache/polaris Apache Polaris is an open-source, fully-featured catalog for Apache Iceberg tables that implements the Iceberg REST catalog API. It provide… | 86 | 2043 | active |
| apache/bookkeeper Apache BookKeeper is a scalable, fault-tolerant, low-latency distributed storage service optimized for append-only workloads, offering repl… | 89 | 2009 | stable |
| Mooncake-Labs/pg_mooncake pg_mooncake is a Postgres extension that maintains a columnstore mirror of Postgres tables in Apache Iceberg, enabling fast real-time analy… | 58 | 2003 | active |
| pravega/pravega Pravega is an open-source distributed storage service that implements Streams as a first-class primitive: durable, elastic, append-only, un… | 27 | 1997 | active |
| elastic/elasticsearch-hadoop Elasticsearch-Hadoop (ES-Hadoop) is a Java connector library that integrates Elasticsearch real-time search and analytics with the Hadoop e… | 95 | 1972 | active |
| fluid-cloudnative/fluid Fluid is a Kubernetes-native distributed dataset orchestrator and accelerator for data-intensive applications such as big data and AI workl… | 79 | 1965 | active |
| apache/cassandra-spark-connector The official Apache Spark connector for Apache Cassandra, letting Spark applications read Cassandra tables as RDDs and DataFrames, write th… | 31 | 1954 | active |
| faust-streaming/faust Faust-streaming is a Python stream processing library that ports Kafka Streams ideas to Python using asyncio, letting developers build dist… | 99 | 1884 | active |
| zhp8341/flink-streaming-platform-web A lightweight web-based management platform built on Apache Flink that lets users configure, deploy, and monitor streaming compute jobs ent… | 48 | 1858 | active |
| pinterest/secor Secor is a Java service that persists Apache Kafka log messages to object stores such as Amazon S3, Google Cloud Storage, Azure Blob Storag… | 64 | 1856 | stable |
| nisshi-io/nisshi Nisshi is a stateless, Kafka API-compatible message broker written in async Rust with pluggable storage engines (PostgreSQL, SQLite/libSQL,… | 88 | 1851 | active |
| alibaba/havenask Havenask is a large-scale distributed information search engine developed by Alibaba Group, written in C++ and used to power search for ser… | 57 | 1836 | active |
| Angel-ML/angel Angel is a high-performance distributed parameter server for large-scale machine learning and graph computing, developed by Tencent and Pek… | 69 | 6787 | maintenance |
| apache/auron Apache Auron (Incubating) is a native vectorized query accelerator for distributed big data engines like Apache Spark, built on Apache Data… | 89 | 1795 | active |
| apache/storm Apache Storm is a free and open-source distributed realtime computation system for processing unbounded streams of data, providing primitiv… | 97 | 6698 | maintenance |
| dbt-labs/metricflow MetricFlow is a Python semantic layer library from dbt Labs that lets you define, build, and maintain business metrics in code. It compiles… | 99 | 1768 | active |
| Netflix/genie Genie is a federated distributed job orchestration and execution engine developed by Netflix, exposing REST APIs to run big data jobs like … | 64 | 1767 | active |
| kairosdb/kairosdb KairosDB is a fast, distributed, scalable time series database written in Java and built on top of Cassandra. It stores and queries metrics… | 54 | 1762 | active |
| TuGraph-family/tugraph-db TuGraph is a high-performance graph database supporting labeled property graphs, full ACID transactions, and OpenCypher queries. It embeds … | 66 | 1760 | active |
| philhagen/sof-elk SOF-ELK is a pre-built virtual appliance based on the Elastic stack (Elasticsearch, Logstash, Kibana, Filebeat) tailored for computer foren… | 76 | 1753 | active |
| dingodb/dingo DingoDB is an open-source distributed multi-modal vector database that combines relational (SQL) and vector semantics in a unified platform… | 64 | 1701 | active |
| matanolabs/matano Matano is an open-source, cloud-native security data lake that runs in your AWS account, normalizing unstructured security logs into a stru… | 23 | 1694 | active |
| apache/solr Apache Solr is a fast, open-source, multi-modal search platform built on Apache Lucene, providing full-text, vector, and geospatial search … | 77 | 1667 | stable |
| FrigadeHQ/trench Trench is an open-source event tracking and analytics infrastructure built on Apache Kafka and ClickHouse, shipped as a single production-r… | 58 | 1661 | active |
| tylertreat/BoomFilters A Go library of probabilistic data structures for processing continuous, unbounded data streams, including Stable, Scalable, Counting, and … | 56 | 1645 | stable |
| almond-sh/almond Almond is a Scala kernel for Jupyter, formerly known as jupyter-scala. It wraps the Ammonite Scala REPL with Jupyter integration, offering … | 84 | 1625 | active |
| Snowflake-Labs/pg_lake pg_lake is a set of PostgreSQL extensions that turn Postgres into a lakehouse engine for Apache Iceberg tables and raw data lake files in o… | 82 | 1625 | active |
| tonbo-io/tonbo Tonbo is an embedded database library written in Rust for serverless and edge runtimes, storing data as Parquet files on S3 with coordinati… | 63 | 1614 | active |
| apache/gluten Apache Gluten is a middle-layer plugin that offloads JVM-based SQL engine execution (primarily Spark SQL) to high-performance native engine… | 91 | 1592 | active |
| getdozer/dozer Dozer is a real-time data movement tool written in Rust that captures change data (CDC) from sources like Postgres, MySQL, Snowflake, and K… | 23 | 1579 | active |
| quixio/quix-streams Quix Streams is a pure Python framework for building real-time data pipelines and event-driven applications on Apache Kafka using a Streami… | 97 | 1568 | active |