Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: big-data

477 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
DataLinkDC/dinky
Dinky is an open-source, one-stop real-time computing platform built around Apache Flink, providing an IDE for FlinkSQL development, online…
803752active
awslabs/deequ
Deequ is a Scala library built on Apache Spark for defining 'unit tests for data' that measure data quality in large datasets. It computes …
933643active
TimelyDataflow/timely-dataflow
A modular Rust implementation of the timely dataflow computational model from the Naiad paper, providing a low-latency cyclic data-parallel…
923642active
apache/arrow-rs
The official Rust implementation of Apache Arrow, providing the in-memory columnar format, and of Apache Parquet for columnar file reading …
983592active
Netflix/atlas
Atlas is Netflix's in-memory dimensional time series database, serving as a backend for storing and querying operational metrics at scale. …
903564active
chrislusf/gleam
Gleam is a fast, efficient distributed map/reduce execution system written in pure Go, defining computations as DAG flows that can run stan…
753562active
holoviz/datashader
Datashader is a Python data rasterization pipeline that renders very large datasets into fixed-size images by projecting, aggregating, and …
893558active
alibaba/GraphScope
GraphScope is a unified distributed graph computing platform from Alibaba that combines graph analytics (GRAPE), interactive graph queries …
733556active
NLPIR-team/NLPIR
NLPIR is a Chinese natural language processing platform (formerly ICTCLAS) providing an SDK with around twenty text analysis functions incl…
753474active
apache/datafusion-sqlparser-rs
An extensible SQL lexer and parser library for Rust that produces an AST conforming to the ANSI/ISO SQL standard and various vendor dialect…
773432active
apache/linkis
Apache Linkis is a computation middleware layer that sits between upper-layer applications and underlying big data engines such as Spark, H…
713406stable
apache/paimon
Apache Paimon is a lakehouse table format that combines lake format with LSM tree structure to support real-time streaming updates. It enab…
873384active
lakehq/sail
Sail is an open-source, Rust-native multimodal compute engine that serves as a drop-in replacement for Apache Spark, unifying batch process…
933333active
apache/avro
Apache Avro is a data serialization system providing rich data structures, a compact binary format, container files, and RPC support, with …
883301stable
root-project/root
ROOT is a C++ data analysis framework developed at CERN for storing, processing, and visualizing large scientific datasets, with a columnar…
963288stable
apache/nutch
Apache Nutch is a highly extensible and scalable open-source web crawler built on Apache Hadoop data structures. It supports batch crawling…
773277stable
WeBankFinTech/DataSphereStudio
DataSphere Studio is a one-stop data application development and management portal from WeBank, built on the Linkis computing middleware. I…
503262active
PeerDB-io/peerdb
PeerDB is an open-source ETL/CDC tool that streams data from Postgres to data warehouses, queues, and storage engines, with an actively mai…
933252active
lakesoul-io/LakeSoul
LakeSoul is a cloud-native, real-time lakehouse framework with a Rust-native core providing ACID table format, concurrent upserts, incremen…
823247active
duckdb/pg_duckdb
pg_duckdb is an official PostgreSQL extension that embeds DuckDB's columnar-vectorized analytics engine inside Postgres. It lets users run …
723210active
apache/gravitino
Apache Gravitino is a high-performance, geo-distributed, federated metadata lake that provides unified metadata management across diverse d…
893188active
vortex-data/vortex
Vortex is an extensible, high-performance columnar file format and toolkit for compressed data, positioned as a faster alternative to Apach…
883158active
apache/hugegraph
Apache HugeGraph is a fast, highly scalable graph database supporting tens of billions of vertices and edges with OLTP real-time queries vi…
733157active
Tencent/Tendis
Tendis is a high-performance distributed key-value storage system fully compatible with the Redis protocol, built on RocksDB for disk-based…
813155active
blockchain-etl/ethereum-etl
A Python CLI tool that extracts Ethereum blockchain data (blocks, transactions, token transfers, logs, traces, contracts) and exports it to…
523127active
alldatacenter/alldata
AllData is a definable data platform (数据中台) that integrates open-source components like DolphinScheduler, DataX, SeaTunnel, OpenMetadata, a…
833082active
apache/parquet-java
Apache Parquet Java is the Java implementation of the Apache Parquet column-oriented data file format. It provides reading and writing of e…
903076active
heavyai/heavydb
HeavyDB (formerly MapD/OmniSciDB) is an open-source SQL-based, relational, columnar database engine that uses CPUs and NVIDIA GPUs to query…
723057active
hydradatabase/columnar
Hydra Columnar is a Postgres extension providing column-oriented table storage via the table access method API, enabling fast analytical qu…
263039stable
TimelyDataflow/differential-dataflow
A Rust library implementing differential dataflow on top of timely dataflow, enabling data-parallel programs that efficiently process large…
932998active
ekzhu/datasketch
datasketch is a Python library of probabilistic data structures (MinHash, Weighted MinHash, HyperLogLog, HyperLogLog++) and indexes (MinHas…
912959stable
duckdb/ducklake
DuckLake is an integrated data lake and lakehouse catalog format that stores data as Parquet files and manages metadata in an ACID-complian…
652939active
yahoo/CMAK
CMAK (Cluster Manager for Apache Kafka, formerly Kafka Manager) is a web-based tool for managing and inspecting Apache Kafka clusters. It s…
2311925maintenance
numaproj/numaflow
Numaflow is a Kubernetes-native, serverless platform for running massively parallel data and streaming processing jobs. It lets developers …
982825active
kafbat/kafka-ui
Kafbat UI is a free, open-source web UI for monitoring and managing Apache Kafka clusters, supporting multi-cluster management, topic and m…
782635active
spotify/scio
Scio is a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. It provides a unified batch and strea…
962628active
OpenLineage/OpenLineage
OpenLineage is an open standard and framework for collecting data lineage metadata, defining a generic model of jobs, runs, and datasets wi…
972626active
meltano/meltano
Meltano is an open-source, declarative, code-first data integration engine and CLI for building and running ELT pipelines. It orchestrates …
972610stable
pingcap/ossinsight
OSSInsight is a web analytics platform that analyzes over 10 billion GitHub events to provide rankings, trends, and comparisons of open sou…
672498active
griddb/griddb
GridDB is an open-source, high-performance distributed database optimized for time-series IoT and big data workloads, offering both NoSQL a…
662475active
confluentinc/schema-registry
Confluent Schema Registry is a Java-based serving layer that stores and retrieves Avro, JSON Schema, and Protobuf schemas via a RESTful API…
772462active
Mycat
MyCAT is an open-source database middleware written in Java that acts as a distributed MySQL cluster proxy, providing sharding, read/write …
609514maintenance
elementary-data/elementary
Elementary OSS is an open-source CLI for dbt-native data observability that reads warehouse metadata, dbt artifacts, and test results to ge…
972398active
apache/sedona
Apache Sedona is a cluster computing framework for processing large-scale geospatial data across Spark, Flink, and Snowflake. It provides S…
902390active
apache/geode
Apache Geode is a distributed in-memory data management platform (originally GemFire) providing real-time, consistent, low-latency access t…
902378stable
quarylabs/quary
Quary is an open-source business intelligence tool for engineers that connects to databases, lets users write SQL to transform and document…
812376active
syslog-ng/syslog-ng
syslog-ng is an enhanced syslog daemon that collects, parses, filters, and routes log messages from a wide range of sources including syslo…
922371stable
moj-analytical-services/splink
Splink is a Python library for fast, scalable probabilistic record linkage and entity resolution, based on the Fellegi-Sunter model with un…
902364active
apache/kyuubi
Apache Kyuubi is a distributed, multi-tenant Thrift JDBC/ODBC gateway that provides serverless SQL access to data warehouses and lakehouses…
942361stable
apache/ambari
Apache Ambari is an open-source tool for provisioning, managing, and monitoring Apache Hadoop clusters. It provides a browser-based managem…
792311active
twitter/algebird
Algebird is a Scala library providing abstract algebra typeclasses (Monoids, Groups, Rings) for building composable aggregation systems. It…
472295stable
pinterest/querybook
Querybook is an open-source big data IDE from Pinterest that combines a collaborative notebook interface (DataDocs) with table metadata fro…
672280active
apache/gobblin
Apache Gobblin is a distributed data integration framework for ingesting, replicating, organizing, and managing lifecycle of data across st…
662270stable
MarquezProject/marquez
Marquez is an open-source metadata service for collecting, aggregating, and visualizing a data ecosystem's metadata. It serves as the refer…
672268active
timeplus-io/proton
Timeplus Proton is a unified streaming SQL engine shipped as a single dependency-free C++ binary, built on the ClickHouse engine. It ingest…
932247active
nathanmarz/storm
Apache Storm (originally nathanmarz/storm) is a distributed, fault-tolerant realtime computation system for stream processing, continuous c…
328765maintenance
ytsaurus/ytsaurus
YTsaurus is an open-source, fault-tolerant big data platform combining distributed storage, MapReduce processing, a SQL query engine, and a…
942200active
vaexio/vaex
Vaex is a high-performance Python DataFrame library for lazy, out-of-core processing of large tabular datasets, using memory mapping and ze…
578507maintenance
reugn/go-streams
A lightweight stream processing library for Go offering a concise DSL to build declarative data pipelines from composable sources, flows, a…
582170active
fugue-project/fugue
Fugue is a Python library providing a unified interface for distributed computing, letting users run Python, Pandas, Polars, and SQL code o…
822169active
tdunning/t-digest
A Java library implementing the t-digest data structure for accurate online accumulation of rank-based statistics such as quantiles and tri…
352166stable
apache/atlas
Apache Atlas is an open-source metadata management and data governance framework providing a shared metadata store, data lineage, and class…
772136active
soedinglab/MMseqs2
MMseqs2 is an open-source C++ suite for ultra-fast, sensitive search and clustering of huge protein and nucleotide sequence sets, running o…
702131active
apache/datafusion-ballista
Apache DataFusion Ballista is a distributed query execution engine built on Apache DataFusion that parallelizes SQL and DataFrame workloads…
982115active
apache/fluss
Apache Fluss is a lakehouse-native streaming storage built for real-time analytics and AI, serving as the real-time data layer for Lakehous…
782113active
warpstreamlabs/bento
Bento is a high-performance, resilient stream processor written in Go that connects a wide range of sources and sinks (Kafka, Pub/Sub, Redi…
912112active
TileDB-Inc/TileDB
TileDB is an embeddable C++ storage engine for dense and sparse multi-dimensional arrays, with cloud object storage support (S3, GCS, Azure…
862073stable
feldera/feldera
Feldera is an incremental computation engine written in Rust that evaluates arbitrary SQL programs incrementally using DBSP theory, maintai…
922064active
zarr-developers/zarr-python
Zarr is a Python library implementing the Zarr storage format for chunked, compressed, N-dimensional arrays with NumPy-compatible dtypes. I…
982044stable
apache/polaris
Apache Polaris is an open-source, fully-featured catalog for Apache Iceberg tables that implements the Iceberg REST catalog API. It provide…
862043active
apache/bookkeeper
Apache BookKeeper is a scalable, fault-tolerant, low-latency distributed storage service optimized for append-only workloads, offering repl…
892009stable
Mooncake-Labs/pg_mooncake
pg_mooncake is a Postgres extension that maintains a columnstore mirror of Postgres tables in Apache Iceberg, enabling fast real-time analy…
582003active
pravega/pravega
Pravega is an open-source distributed storage service that implements Streams as a first-class primitive: durable, elastic, append-only, un…
271997active
elastic/elasticsearch-hadoop
Elasticsearch-Hadoop (ES-Hadoop) is a Java connector library that integrates Elasticsearch real-time search and analytics with the Hadoop e…
951972active
fluid-cloudnative/fluid
Fluid is a Kubernetes-native distributed dataset orchestrator and accelerator for data-intensive applications such as big data and AI workl…
791965active
apache/cassandra-spark-connector
The official Apache Spark connector for Apache Cassandra, letting Spark applications read Cassandra tables as RDDs and DataFrames, write th…
311954active
faust-streaming/faust
Faust-streaming is a Python stream processing library that ports Kafka Streams ideas to Python using asyncio, letting developers build dist…
991884active
zhp8341/flink-streaming-platform-web
A lightweight web-based management platform built on Apache Flink that lets users configure, deploy, and monitor streaming compute jobs ent…
481858active
pinterest/secor
Secor is a Java service that persists Apache Kafka log messages to object stores such as Amazon S3, Google Cloud Storage, Azure Blob Storag…
641856stable
nisshi-io/nisshi
Nisshi is a stateless, Kafka API-compatible message broker written in async Rust with pluggable storage engines (PostgreSQL, SQLite/libSQL,…
881851active
alibaba/havenask
Havenask is a large-scale distributed information search engine developed by Alibaba Group, written in C++ and used to power search for ser…
571836active
Angel-ML/angel
Angel is a high-performance distributed parameter server for large-scale machine learning and graph computing, developed by Tencent and Pek…
696787maintenance
apache/auron
Apache Auron (Incubating) is a native vectorized query accelerator for distributed big data engines like Apache Spark, built on Apache Data…
891795active
apache/storm
Apache Storm is a free and open-source distributed realtime computation system for processing unbounded streams of data, providing primitiv…
976698maintenance
dbt-labs/metricflow
MetricFlow is a Python semantic layer library from dbt Labs that lets you define, build, and maintain business metrics in code. It compiles…
991768active
Netflix/genie
Genie is a federated distributed job orchestration and execution engine developed by Netflix, exposing REST APIs to run big data jobs like …
641767active
kairosdb/kairosdb
KairosDB is a fast, distributed, scalable time series database written in Java and built on top of Cassandra. It stores and queries metrics…
541762active
TuGraph-family/tugraph-db
TuGraph is a high-performance graph database supporting labeled property graphs, full ACID transactions, and OpenCypher queries. It embeds …
661760active
philhagen/sof-elk
SOF-ELK is a pre-built virtual appliance based on the Elastic stack (Elasticsearch, Logstash, Kibana, Filebeat) tailored for computer foren…
761753active
dingodb/dingo
DingoDB is an open-source distributed multi-modal vector database that combines relational (SQL) and vector semantics in a unified platform…
641701active
matanolabs/matano
Matano is an open-source, cloud-native security data lake that runs in your AWS account, normalizing unstructured security logs into a stru…
231694active
apache/solr
Apache Solr is a fast, open-source, multi-modal search platform built on Apache Lucene, providing full-text, vector, and geospatial search …
771667stable
FrigadeHQ/trench
Trench is an open-source event tracking and analytics infrastructure built on Apache Kafka and ClickHouse, shipped as a single production-r…
581661active
tylertreat/BoomFilters
A Go library of probabilistic data structures for processing continuous, unbounded data streams, including Stable, Scalable, Counting, and …
561645stable
almond-sh/almond
Almond is a Scala kernel for Jupyter, formerly known as jupyter-scala. It wraps the Ammonite Scala REPL with Jupyter integration, offering …
841625active
Snowflake-Labs/pg_lake
pg_lake is a set of PostgreSQL extensions that turn Postgres into a lakehouse engine for Apache Iceberg tables and raw data lake files in o…
821625active
tonbo-io/tonbo
Tonbo is an embedded database library written in Rust for serverless and edge runtimes, storing data as Parquet files on S3 with coordinati…
631614active
apache/gluten
Apache Gluten is a middle-layer plugin that offloads JVM-based SQL engine execution (primarily Spark SQL) to high-performance native engine…
911592active
getdozer/dozer
Dozer is a real-time data movement tool written in Rust that captures change data (CDC) from sources like Postgres, MySQL, Snowflake, and K…
231579active
quixio/quix-streams
Quix Streams is a pure Python framework for building real-time data pipelines and event-driven applications on Apache Kafka using a Streami…
971568active

← prev page 2 / 5 next →