domain: big-data
477 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| holdenk/spark-testing-base A library providing base classes for writing tests for Apache Spark applications in Scala and Python. It handles the setup and teardown of … | 66 | 1555 | active |
| substrait-io/substrait Substrait is a cross-language specification and tooling for describing data compute operations as standardized relational algebra query pla… | 99 | 1549 | active |
| combust/mleap MLeap is a serialization format (Bundle.ML) and portable execution engine for machine learning pipelines, implemented in Scala with Python … | 95 | 1543 | active |
| google/tensorstore TensorStore is a C++ and Python library for reading and writing large multi-dimensional arrays with a uniform API across formats like zarr,… | 77 | 1536 | active |
| hi-primus/optimus Optimus is a Python library for agile data preparation that provides a unified API over pandas, Dask, cuDF, Dask-cuDF, Vaex, and PySpark. I… | 23 | 1536 | active |
| BemiHQ/BemiDB BemiDB is an open-source analytical data warehouse that combines built-in data source connectors (like Fivetran) with a Postgres-compatible… | 55 | 1531 | active |
| projectnessie/nessie Project Nessie is a transactional catalog for data lakes that provides Git-like branching, tagging, and cross-table transactions for Apache… | 98 | 1497 | active |
| apache/inlong Apache InLong is a one-stop, full-scenario integration framework for massive data, supporting data ingestion, synchronization, and subscrip… | 90 | 1497 | stable |
| locationtech/geomesa GeoMesa is an open-source suite of JVM-based tools for large-scale geospatial querying and analytics on distributed computing systems. It p… | 78 | 1494 | active |
| dremio/dremio-oss Dremio is an open-source data lakehouse platform that lets users query and analyze data across data lakes, object storage, and databases wi… | 52 | 1491 | active |
| Netflix/mantis Mantis is a Netflix open-source stream processing platform written in Java for building realtime, cost-effective, operations-focused applic… | 67 | 1470 | active |
| apache/hop Apache Hop is an open-source data and metadata orchestration platform for visually designing and running data integration pipelines and wor… | 94 | 1447 | active |
| datazip-inc/olake OLake Go is a high-performance open-source EL (extract-load) engine written in Go that replicates databases (PostgreSQL, MySQL, MongoDB, Or… | 84 | 1431 | active |
| xitongsys/parquet-go A pure-Go library for reading and writing Apache Parquet files, supporting nested and flat schemas with multiple encodings. It maps Go stru… | 59 | 1428 | active |
| lakekeeper/lakekeeper Lakekeeper is an Apache Iceberg REST Catalog implementation written in Rust that provides centralized access control, credential vending, a… | 91 | 1426 | active |
| wgzhao/Addax Addax is a fast, extensible ETL tool for synchronizing data between heterogeneous SQL and NoSQL data sources, forked and evolved from Aliba… | 98 | 1425 | active |
| opendatadiscovery/odd-platform ODD Platform is an open-source data discovery and observability platform that provides a federated data catalog, end-to-end data and micros… | 92 | 1425 | active |
| cloudera/hue Hue is an open-source web-based SQL query assistant and editor for databases and data warehouses, supporting connectors like Hive, Impala, … | 67 | 1409 | active |
| OpenTSDB/opentsdb OpenTSDB is a distributed, scalable time series database (TSDB) built on top of HBase, written in Java. It stores, indexes, and serves metr… | 23 | 5065 | maintenance |
| heterodb/pg-strom PG-Strom is a PostgreSQL extension that accelerates SQL analytics and batch workloads using GPU devices, NVMe-SSD storage, and Apache Arrow… | 76 | 1408 | active |
| apache/iceberg-rust A Rust implementation of the Apache Iceberg open table format for managing large analytic datasets. It provides crates for the core Iceberg… | 90 | 1390 | active |
| mtth/avsc A pure JavaScript implementation of the Apache Avro serialization specification for Node.js and browsers. It provides fast, compact binary … | 53 | 1384 | stable |
| aerospike/aerospike-server Aerospike is a distributed, flash-optimized NoSQL database server written in C that combines in-memory speeds with SSD-based durability thr… | 90 | 1375 | stable |
| apache/cloudberry Apache Cloudberry is an advanced open-source Massively Parallel Processing (MPP) database derived from Greenplum and built on a modern Post… | 78 | 1374 | active |
| locationtech/geotrellis GeoTrellis is a Scala library and framework for high-performance reading, writing, and processing of geospatial raster and vector data. It … | 84 | 1372 | stable |
| PyTables/PyTables PyTables is a Python package for managing hierarchical datasets built on top of the HDF5 library and NumPy. It provides a fast, object-orie… | 79 | 1372 | stable |
| jupyter-incubator/sparkmagic Sparkmagic is a set of Jupyter magics and kernels for interactively working with remote Spark clusters through a REST server such as Livy, … | 41 | 1366 | active |
| XTXMarkets/ternfs TernFS is an exabyte-scale, multi-region distributed file system written in C++, designed for storing large immutable files on commodity ha… | 62 | 1339 | active |
| GoogleCloudPlatform/DataflowTemplates A collection of Google-provided Apache Beam pipeline templates for Google Cloud Dataflow that solve common in-cloud data tasks like import/… | 95 | 1310 | active |
| GoogleCloudPlatform/bigquery-utils A Google-maintained collection of utilities for BigQuery, including SQL user-defined functions, stored procedures, views, Python and shell … | 73 | 1306 | active |
| arkflow-rs/arkflow ArkFlow is a high-performance stream processing engine written in Rust on top of Tokio, connecting configurable inputs (Kafka, MQTT, HTTP, … | 70 | 1302 | active |
| apache/impala Apache Impala is a massively parallel, distributed SQL query engine written in C++ for analyzing petabyte-scale data stored in open data an… | 67 | 1285 | stable |
| DTStack/Taier Taier is a self-hosted distributed dispatching platform for big data that handles task submission, DAG-based scheduling, and operations/mai… | 84 | 1284 | active |
| azkaban/azkaban Azkaban is an open-source batch workflow job scheduler created at LinkedIn to run Hadoop and data warehouse jobs. It resolves job execution… | 23 | 4508 | maintenance |
| apache/ozone Apache Ozone is a scalable, redundant, distributed object store for Hadoop and cloud-native environments, supporting billions of objects wi… | 95 | 1271 | stable |
| water8394/flink-recommandSystem-demo A real-time product recommendation system built on Apache Flink, demonstrating streaming computation of product popularity, user profiles, … | 32 | 4480 | maintenance |
| apache/datafusion-comet Apache DataFusion Comet is a high-performance accelerator plugin for Apache Spark that keeps queries Arrow-native end-to-end, executing ope… | 81 | 1262 | active |
| mesos/chronos Chronos is a distributed, fault-tolerant job scheduler that runs on Apache Mesos as a replacement for cron. It supports ISO8601 repeating i… | 23 | 4376 | maintenance |
| zinggAI/zingg Zingg is an ML-based tool for scalable master data management, entity resolution, identity resolution, and record deduplication. It runs on… | 88 | 1243 | active |
| baidu/BaikalDB BaikalDB is a distributed HTAP (hybrid transactional/analytical processing) database developed by Baidu, written in C++ and using Raft for … | 82 | 1238 | active |
| mukunku/ParquetViewer A free Windows desktop application for viewing and querying Apache Parquet files. It supports SQL queries, metadata inspection, partitioned… | 92 | 1228 | active |
| citusdata/postgresql-hll A PostgreSQL extension that adds HyperLogLog as a native data type for approximate distinct value counting with tunable precision. It uses … | 83 | 1228 | stable |
| kevwan/go-stash go-stash is a high-performance, open-source server-side data processing pipeline written in Go that ingests data from Kafka, applies config… | 60 | 1224 | active |
| skyplane-project/skyplane Skyplane is a CLI tool for blazingly fast bulk data transfers between cloud object stores (AWS S3, Azure Blob, GCS, IBM COS) and local disk… | 23 | 1216 | active |
| apache/incubator-xtable Apache XTable (incubating) is a cross-table converter that translates lakehouse table format metadata between Apache Hudi, Apache Iceberg, … | 84 | 1205 | active |
| graphframes/graphframes GraphFrames is a package for Apache Spark that provides DataFrame-based graph processing with distributed graph algorithms like PageRank, c… | 95 | 1203 | stable |
| gityuanbao/share A personal open-source repository whose main component is akshare_collector, a Python tool built on AKShare that collects Chinese financial… | 64 | 1201 | active |
| xorbitsai/xorbits Xorbits is an open-source distributed computing framework that scales Python data science and machine learning workloads from a laptop to l… | 64 | 1199 | active |
| marsupialtail/quokka Quokka is a lightweight distributed dataflow/query engine written in Python, built on Ray, DuckDB, Polars, and Arrow, designed for stateful… | 23 | 1192 | active |
| pentaho/mondrian Mondrian is an open-source OLAP (Online Analytical Processing) server written in Java that lets business users analyze large volumes of rel… | 77 | 1174 | active |
| apache/amoro Apache Amoro (incubating) is a Lakehouse management system built on open data lake formats like Iceberg, Paimon, and Mixed-Hive. It provide… | 73 | 1171 | active |
| apache/accumulo Apache Accumulo is a sorted, distributed key/value store built on Apache Hadoop's HDFS and Apache ZooKeeper for scalable storage and retrie… | 77 | 1167 | stable |
| Riak Riak is a decentralized, distributed NoSQL key-value datastore built in Erlang by Basho Technologies, designed for high availability and op… | 66 | 4027 | maintenance |
| Edgio/vflow vFlow is a high-performance, scalable network flow collector written in pure Go that ingests IPFIX (RFC7011), sFlow v5, and Netflow v5/v9 t… | 23 | 1155 | active |
| baidu/bigflow Baidu Bigflow is a distributed computing framework offering simple, flexible Python APIs for writing data processing programs that can run … | 48 | 1131 | active |
| apache/iceberg-python PyIceberg is a Python implementation of the Apache Iceberg table format specification, providing programmatic access to Iceberg table metad… | 92 | 1124 | active |
| scratchdata/scratchdata Scratch Data is a self-hosted Go service that acts as a wrapper around analytics databases like DuckDB, ClickHouse, BigQuery, Snowflake, an… | 19 | 1120 | active |
| yahoo/TensorFlowOnSpark TensorFlowOnSpark is a Python library that lets existing TensorFlow programs run distributed training and inference on Apache Spark and Had… | 23 | 3845 | maintenance |
| childe/gohangout Gohangout is a Logstash-like data processing tool written in Go that consumes events (commonly from Kafka), applies configurable filters, a… | 74 | 1103 | active |
| apache/systemds Apache SystemDS is an open-source machine learning system covering the end-to-end data science lifecycle, from data cleaning and feature en… | 87 | 1097 | active |
| y-scope/clp CLP (Compressed Log Processor) is an open-source log management tool that losslessly compresses JSON and unstructured text logs and enables… | 94 | 1083 | active |
| apache/ranger Apache Ranger is a framework to enable, monitor, and manage comprehensive data security across the Hadoop platform and beyond. It provides … | 77 | 1074 | stable |
| linkedin/databus Databus is LinkedIn's source-agnostic distributed change data capture (CDC) system that reliably captures and streams primary data store ch… | 32 | 3679 | maintenance |
| hail-is/hail Hail is an open-source Python library for scalable exploration and analysis of genomic data, built on Spark, Scala, and C++ primitives for … | 91 | 1070 | active |
| lensesio/stream-reactor Stream Reactor is a collection of Apache 2.0 licensed Kafka Connect sinks and sources maintained by Lenses.io since 2016. It provides conne… | 95 | 1068 | active |
| apache/celeborn Apache Celeborn is an elastic, high-performance intermediate data service for big data compute engines, focused on managing shuffle and spi… | 95 | 1061 | active |
| apache/phoenix Apache Phoenix is a SQL layer over Apache HBase delivered as a client-embedded JDBC driver, compiling standard SQL queries into low-latency… | 88 | 1060 | active |
| sirius-db/sirius Sirius is a GPU-native SQL analytics engine written in C++ that accelerates query execution by offloading it to GPUs. It integrates with ex… | 69 | 1059 | active |
| bigdatagenomics/adam ADAM is a genomics analysis platform built on Apache Spark that provides schemas and APIs for processing genomic data like reads, variants,… | 55 | 1057 | active |
| alibaba/Alink Alink is a machine learning algorithm platform built on Apache Flink, developed by Alibaba's PAI team. It provides a large library of batch… | 23 | 3611 | maintenance |
| devlive-community/datacap DataCap is a self-hosted integrated platform for managing, transforming, integrating, and visualizing data across many data sources. It pro… | 73 | 1051 | active |
| axiomhq/hyperloglog A Go library implementing the HyperLogLog algorithm for approximating the number of distinct elements in a multiset, using LogLog-Beta bias… | 84 | 1047 | active |
| airbnb/chronon Chronon is an open-source data platform from Airbnb for computing, backfilling, and serving ML features. It handles batch and streaming fea… | 75 | 1043 | active |
| sqlparser/sqlflow_public SQLFlow is a tool that tracks column-level data lineage from SQL scripts across more than 20 major databases like Snowflake, Hive, Oracle, … | 69 | 1041 | active |
| myscale/MyScaleDB MyScaleDB is a SQL vector database built as a fork of ClickHouse, adding high-performance vector search and full-text search to a proven OL… | 25 | 1039 | active |
| twitter/scalding Scalding is a Scala library for specifying Hadoop MapReduce jobs, built on top of the Cascading framework. It provides a type-safe, functio… | 23 | 3523 | maintenance |
| apache/flink-kubernetes-operator A Kubernetes operator for Apache Flink, implemented in Java, that manages the full lifecycle of Flink Application, Session, and Job deploym… | 76 | 1027 | stable |
| apache/yunikorn-core Apache YuniKorn Core is the core scheduling engine of a lightweight, universal resource scheduler for container orchestrators, primarily de… | 90 | 1025 | active |
| liyupi/sql-generator A web application that generates structured, reusable SQL statements from JSON definitions, built with Vue3, TypeScript, Vite, Ant Design, … | 32 | 3426 | maintenance |
| Logflare/logflare Logflare is a centralized structured log ingestion and querying service that streams log events into BigQuery (or its own backend) with aut… | 96 | 1004 | active |
| databricks/koalas Koalas implements the pandas DataFrame API on top of Apache Spark, letting data scientists use familiar pandas code on distributed big data… | 23 | 3372 | maintenance |
| chrislusf/glow Glow is a pure-Go distributed computation library providing MapReduce-style flow APIs (Map, Filter, Reduce) that can run in parallel on a s… | 32 | 3218 | maintenance |
| blaze/blaze Blaze is a Python library that provides a NumPy/Pandas-like interface for querying data living in databases, files, and other computing sys… | 23 | 3189 | maintenance |
| spark-notebook/spark-notebook An open-source web-based notebook for interactive and reactive data science using Scala and Apache Spark. It combines Scala code, SQL, mark… | 23 | 3140 | maintenance |
| TuiQiao/CBoard CBoard is a self-service open-source business intelligence platform for designing interactive multi-dimensional OLAP reports and dashboards… | 57 | 3097 | maintenance |
| uber/aresdb AresDB is a GPU-powered real-time analytics storage and query engine written in Go and C++ with CUDA kernels. It provides low query latency… | 23 | 3078 | maintenance |
| CodeRayZhang/Movie_Recommend A full-stack movie recommendation system built on Spark, including a Scrapy crawler, an SSM-based movie website, an admin backend, and a Sp… | 32 | 3007 | maintenance |
| baidu/bfs Baidu File System (BFS) is a distributed file system written in C++ designed for real-time applications, offering fault tolerance, low read… | 23 | 2847 | maintenance |
| spark-jobserver/spark-jobserver Spark Job Server is a RESTful service for submitting, managing, and monitoring Apache Spark jobs, jars, and job contexts. It supports persi… | 54 | 2836 | maintenance |
| TalkingData/inmap inMap is a JavaScript big-data geographic visualization library built on top of Baidu Maps. It renders large point datasets as scatter plot… | 23 | 2809 | maintenance |
| mars-project/mars Mars is a tensor-based unified framework for large-scale data computation that provides NumPy-, pandas-, and scikit-learn-compatible APIs w… | 23 | 2742 | maintenance |
| douban/dpark DPark is a Python clone of Apache Spark, providing a MapReduce-like distributed computing framework that supports iterative computation. Jo… | 10 | 2663 | maintenance |
| pipelinedb/pipelinedb PipelineDB is a PostgreSQL extension for high-performance time-series aggregation, letting you define continuous SQL queries that increment… | 23 | 2662 | maintenance |
| steveloughran/winutils Windows native binaries (winutils.exe and Hadoop DLLs) for various Hadoop versions, built from the same git commits as official ASF release… | 23 | 2639 | maintenance |
| Yelp/mrjob mrjob is a Python library for writing and running Hadoop Streaming MapReduce jobs, with support for Amazon EMR, Google Cloud Dataproc, self… | 66 | 2613 | maintenance |
| apache/logging-flume Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log-lik… | 77 | 2567 | maintenance |
| justinzm/gopup GoPUP is a Python library that provides convenient interfaces to a wide range of public Chinese data sources, including Baidu/Weibo/Google … | 23 | 2548 | maintenance |
| geekyouth/SZT-bigdata A big data passenger flow analysis system for the Shenzhen Metro, built on Shenzhen Tong smart-card swipe data. It demonstrates ETL and ana… | 60 | 2475 | maintenance |
| cdarlint/winutils A repository of precompiled Windows binaries (winutils.exe, hadoop.dll, hdfs.dll) needed to run Hadoop natively on Windows. It continues th… | 32 | 2287 | maintenance |
| salesforce/TransmogrifAI TransmogrifAI is an AutoML library written in Scala that runs on Apache Spark for building modular, reusable, strongly typed machine learni… | 61 | 2277 | maintenance |