Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: big-data

477 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
holdenk/spark-testing-base
A library providing base classes for writing tests for Apache Spark applications in Scala and Python. It handles the setup and teardown of …
661555active
substrait-io/substrait
Substrait is a cross-language specification and tooling for describing data compute operations as standardized relational algebra query pla…
991549active
combust/mleap
MLeap is a serialization format (Bundle.ML) and portable execution engine for machine learning pipelines, implemented in Scala with Python …
951543active
google/tensorstore
TensorStore is a C++ and Python library for reading and writing large multi-dimensional arrays with a uniform API across formats like zarr,…
771536active
hi-primus/optimus
Optimus is a Python library for agile data preparation that provides a unified API over pandas, Dask, cuDF, Dask-cuDF, Vaex, and PySpark. I…
231536active
BemiHQ/BemiDB
BemiDB is an open-source analytical data warehouse that combines built-in data source connectors (like Fivetran) with a Postgres-compatible…
551531active
projectnessie/nessie
Project Nessie is a transactional catalog for data lakes that provides Git-like branching, tagging, and cross-table transactions for Apache…
981497active
apache/inlong
Apache InLong is a one-stop, full-scenario integration framework for massive data, supporting data ingestion, synchronization, and subscrip…
901497stable
locationtech/geomesa
GeoMesa is an open-source suite of JVM-based tools for large-scale geospatial querying and analytics on distributed computing systems. It p…
781494active
dremio/dremio-oss
Dremio is an open-source data lakehouse platform that lets users query and analyze data across data lakes, object storage, and databases wi…
521491active
Netflix/mantis
Mantis is a Netflix open-source stream processing platform written in Java for building realtime, cost-effective, operations-focused applic…
671470active
apache/hop
Apache Hop is an open-source data and metadata orchestration platform for visually designing and running data integration pipelines and wor…
941447active
datazip-inc/olake
OLake Go is a high-performance open-source EL (extract-load) engine written in Go that replicates databases (PostgreSQL, MySQL, MongoDB, Or…
841431active
xitongsys/parquet-go
A pure-Go library for reading and writing Apache Parquet files, supporting nested and flat schemas with multiple encodings. It maps Go stru…
591428active
lakekeeper/lakekeeper
Lakekeeper is an Apache Iceberg REST Catalog implementation written in Rust that provides centralized access control, credential vending, a…
911426active
wgzhao/Addax
Addax is a fast, extensible ETL tool for synchronizing data between heterogeneous SQL and NoSQL data sources, forked and evolved from Aliba…
981425active
opendatadiscovery/odd-platform
ODD Platform is an open-source data discovery and observability platform that provides a federated data catalog, end-to-end data and micros…
921425active
cloudera/hue
Hue is an open-source web-based SQL query assistant and editor for databases and data warehouses, supporting connectors like Hive, Impala, …
671409active
OpenTSDB/opentsdb
OpenTSDB is a distributed, scalable time series database (TSDB) built on top of HBase, written in Java. It stores, indexes, and serves metr…
235065maintenance
heterodb/pg-strom
PG-Strom is a PostgreSQL extension that accelerates SQL analytics and batch workloads using GPU devices, NVMe-SSD storage, and Apache Arrow…
761408active
apache/iceberg-rust
A Rust implementation of the Apache Iceberg open table format for managing large analytic datasets. It provides crates for the core Iceberg…
901390active
mtth/avsc
A pure JavaScript implementation of the Apache Avro serialization specification for Node.js and browsers. It provides fast, compact binary …
531384stable
aerospike/aerospike-server
Aerospike is a distributed, flash-optimized NoSQL database server written in C that combines in-memory speeds with SSD-based durability thr…
901375stable
apache/cloudberry
Apache Cloudberry is an advanced open-source Massively Parallel Processing (MPP) database derived from Greenplum and built on a modern Post…
781374active
locationtech/geotrellis
GeoTrellis is a Scala library and framework for high-performance reading, writing, and processing of geospatial raster and vector data. It …
841372stable
PyTables/PyTables
PyTables is a Python package for managing hierarchical datasets built on top of the HDF5 library and NumPy. It provides a fast, object-orie…
791372stable
jupyter-incubator/sparkmagic
Sparkmagic is a set of Jupyter magics and kernels for interactively working with remote Spark clusters through a REST server such as Livy, …
411366active
XTXMarkets/ternfs
TernFS is an exabyte-scale, multi-region distributed file system written in C++, designed for storing large immutable files on commodity ha…
621339active
GoogleCloudPlatform/DataflowTemplates
A collection of Google-provided Apache Beam pipeline templates for Google Cloud Dataflow that solve common in-cloud data tasks like import/…
951310active
GoogleCloudPlatform/bigquery-utils
A Google-maintained collection of utilities for BigQuery, including SQL user-defined functions, stored procedures, views, Python and shell …
731306active
arkflow-rs/arkflow
ArkFlow is a high-performance stream processing engine written in Rust on top of Tokio, connecting configurable inputs (Kafka, MQTT, HTTP, …
701302active
apache/impala
Apache Impala is a massively parallel, distributed SQL query engine written in C++ for analyzing petabyte-scale data stored in open data an…
671285stable
DTStack/Taier
Taier is a self-hosted distributed dispatching platform for big data that handles task submission, DAG-based scheduling, and operations/mai…
841284active
azkaban/azkaban
Azkaban is an open-source batch workflow job scheduler created at LinkedIn to run Hadoop and data warehouse jobs. It resolves job execution…
234508maintenance
apache/ozone
Apache Ozone is a scalable, redundant, distributed object store for Hadoop and cloud-native environments, supporting billions of objects wi…
951271stable
water8394/flink-recommandSystem-demo
A real-time product recommendation system built on Apache Flink, demonstrating streaming computation of product popularity, user profiles, …
324480maintenance
apache/datafusion-comet
Apache DataFusion Comet is a high-performance accelerator plugin for Apache Spark that keeps queries Arrow-native end-to-end, executing ope…
811262active
mesos/chronos
Chronos is a distributed, fault-tolerant job scheduler that runs on Apache Mesos as a replacement for cron. It supports ISO8601 repeating i…
234376maintenance
zinggAI/zingg
Zingg is an ML-based tool for scalable master data management, entity resolution, identity resolution, and record deduplication. It runs on…
881243active
baidu/BaikalDB
BaikalDB is a distributed HTAP (hybrid transactional/analytical processing) database developed by Baidu, written in C++ and using Raft for …
821238active
mukunku/ParquetViewer
A free Windows desktop application for viewing and querying Apache Parquet files. It supports SQL queries, metadata inspection, partitioned…
921228active
citusdata/postgresql-hll
A PostgreSQL extension that adds HyperLogLog as a native data type for approximate distinct value counting with tunable precision. It uses …
831228stable
kevwan/go-stash
go-stash is a high-performance, open-source server-side data processing pipeline written in Go that ingests data from Kafka, applies config…
601224active
skyplane-project/skyplane
Skyplane is a CLI tool for blazingly fast bulk data transfers between cloud object stores (AWS S3, Azure Blob, GCS, IBM COS) and local disk…
231216active
apache/incubator-xtable
Apache XTable (incubating) is a cross-table converter that translates lakehouse table format metadata between Apache Hudi, Apache Iceberg, …
841205active
graphframes/graphframes
GraphFrames is a package for Apache Spark that provides DataFrame-based graph processing with distributed graph algorithms like PageRank, c…
951203stable
gityuanbao/share
A personal open-source repository whose main component is akshare_collector, a Python tool built on AKShare that collects Chinese financial…
641201active
xorbitsai/xorbits
Xorbits is an open-source distributed computing framework that scales Python data science and machine learning workloads from a laptop to l…
641199active
marsupialtail/quokka
Quokka is a lightweight distributed dataflow/query engine written in Python, built on Ray, DuckDB, Polars, and Arrow, designed for stateful…
231192active
pentaho/mondrian
Mondrian is an open-source OLAP (Online Analytical Processing) server written in Java that lets business users analyze large volumes of rel…
771174active
apache/amoro
Apache Amoro (incubating) is a Lakehouse management system built on open data lake formats like Iceberg, Paimon, and Mixed-Hive. It provide…
731171active
apache/accumulo
Apache Accumulo is a sorted, distributed key/value store built on Apache Hadoop's HDFS and Apache ZooKeeper for scalable storage and retrie…
771167stable
Riak
Riak is a decentralized, distributed NoSQL key-value datastore built in Erlang by Basho Technologies, designed for high availability and op…
664027maintenance
Edgio/vflow
vFlow is a high-performance, scalable network flow collector written in pure Go that ingests IPFIX (RFC7011), sFlow v5, and Netflow v5/v9 t…
231155active
baidu/bigflow
Baidu Bigflow is a distributed computing framework offering simple, flexible Python APIs for writing data processing programs that can run …
481131active
apache/iceberg-python
PyIceberg is a Python implementation of the Apache Iceberg table format specification, providing programmatic access to Iceberg table metad…
921124active
scratchdata/scratchdata
Scratch Data is a self-hosted Go service that acts as a wrapper around analytics databases like DuckDB, ClickHouse, BigQuery, Snowflake, an…
191120active
yahoo/TensorFlowOnSpark
TensorFlowOnSpark is a Python library that lets existing TensorFlow programs run distributed training and inference on Apache Spark and Had…
233845maintenance
childe/gohangout
Gohangout is a Logstash-like data processing tool written in Go that consumes events (commonly from Kafka), applies configurable filters, a…
741103active
apache/systemds
Apache SystemDS is an open-source machine learning system covering the end-to-end data science lifecycle, from data cleaning and feature en…
871097active
y-scope/clp
CLP (Compressed Log Processor) is an open-source log management tool that losslessly compresses JSON and unstructured text logs and enables…
941083active
apache/ranger
Apache Ranger is a framework to enable, monitor, and manage comprehensive data security across the Hadoop platform and beyond. It provides …
771074stable
linkedin/databus
Databus is LinkedIn's source-agnostic distributed change data capture (CDC) system that reliably captures and streams primary data store ch…
323679maintenance
hail-is/hail
Hail is an open-source Python library for scalable exploration and analysis of genomic data, built on Spark, Scala, and C++ primitives for …
911070active
lensesio/stream-reactor
Stream Reactor is a collection of Apache 2.0 licensed Kafka Connect sinks and sources maintained by Lenses.io since 2016. It provides conne…
951068active
apache/celeborn
Apache Celeborn is an elastic, high-performance intermediate data service for big data compute engines, focused on managing shuffle and spi…
951061active
apache/phoenix
Apache Phoenix is a SQL layer over Apache HBase delivered as a client-embedded JDBC driver, compiling standard SQL queries into low-latency…
881060active
sirius-db/sirius
Sirius is a GPU-native SQL analytics engine written in C++ that accelerates query execution by offloading it to GPUs. It integrates with ex…
691059active
bigdatagenomics/adam
ADAM is a genomics analysis platform built on Apache Spark that provides schemas and APIs for processing genomic data like reads, variants,…
551057active
alibaba/Alink
Alink is a machine learning algorithm platform built on Apache Flink, developed by Alibaba's PAI team. It provides a large library of batch…
233611maintenance
devlive-community/datacap
DataCap is a self-hosted integrated platform for managing, transforming, integrating, and visualizing data across many data sources. It pro…
731051active
axiomhq/hyperloglog
A Go library implementing the HyperLogLog algorithm for approximating the number of distinct elements in a multiset, using LogLog-Beta bias…
841047active
airbnb/chronon
Chronon is an open-source data platform from Airbnb for computing, backfilling, and serving ML features. It handles batch and streaming fea…
751043active
sqlparser/sqlflow_public
SQLFlow is a tool that tracks column-level data lineage from SQL scripts across more than 20 major databases like Snowflake, Hive, Oracle, …
691041active
myscale/MyScaleDB
MyScaleDB is a SQL vector database built as a fork of ClickHouse, adding high-performance vector search and full-text search to a proven OL…
251039active
twitter/scalding
Scalding is a Scala library for specifying Hadoop MapReduce jobs, built on top of the Cascading framework. It provides a type-safe, functio…
233523maintenance
apache/flink-kubernetes-operator
A Kubernetes operator for Apache Flink, implemented in Java, that manages the full lifecycle of Flink Application, Session, and Job deploym…
761027stable
apache/yunikorn-core
Apache YuniKorn Core is the core scheduling engine of a lightweight, universal resource scheduler for container orchestrators, primarily de…
901025active
liyupi/sql-generator
A web application that generates structured, reusable SQL statements from JSON definitions, built with Vue3, TypeScript, Vite, Ant Design, …
323426maintenance
Logflare/logflare
Logflare is a centralized structured log ingestion and querying service that streams log events into BigQuery (or its own backend) with aut…
961004active
databricks/koalas
Koalas implements the pandas DataFrame API on top of Apache Spark, letting data scientists use familiar pandas code on distributed big data…
233372maintenance
chrislusf/glow
Glow is a pure-Go distributed computation library providing MapReduce-style flow APIs (Map, Filter, Reduce) that can run in parallel on a s…
323218maintenance
blaze/blaze
Blaze is a Python library that provides a NumPy/Pandas-like interface for querying data living in databases, files, and other computing sys…
233189maintenance
spark-notebook/spark-notebook
An open-source web-based notebook for interactive and reactive data science using Scala and Apache Spark. It combines Scala code, SQL, mark…
233140maintenance
TuiQiao/CBoard
CBoard is a self-service open-source business intelligence platform for designing interactive multi-dimensional OLAP reports and dashboards…
573097maintenance
uber/aresdb
AresDB is a GPU-powered real-time analytics storage and query engine written in Go and C++ with CUDA kernels. It provides low query latency…
233078maintenance
CodeRayZhang/Movie_Recommend
A full-stack movie recommendation system built on Spark, including a Scrapy crawler, an SSM-based movie website, an admin backend, and a Sp…
323007maintenance
baidu/bfs
Baidu File System (BFS) is a distributed file system written in C++ designed for real-time applications, offering fault tolerance, low read…
232847maintenance
spark-jobserver/spark-jobserver
Spark Job Server is a RESTful service for submitting, managing, and monitoring Apache Spark jobs, jars, and job contexts. It supports persi…
542836maintenance
TalkingData/inmap
inMap is a JavaScript big-data geographic visualization library built on top of Baidu Maps. It renders large point datasets as scatter plot…
232809maintenance
mars-project/mars
Mars is a tensor-based unified framework for large-scale data computation that provides NumPy-, pandas-, and scikit-learn-compatible APIs w…
232742maintenance
douban/dpark
DPark is a Python clone of Apache Spark, providing a MapReduce-like distributed computing framework that supports iterative computation. Jo…
102663maintenance
pipelinedb/pipelinedb
PipelineDB is a PostgreSQL extension for high-performance time-series aggregation, letting you define continuous SQL queries that increment…
232662maintenance
steveloughran/winutils
Windows native binaries (winutils.exe and Hadoop DLLs) for various Hadoop versions, built from the same git commits as official ASF release…
232639maintenance
Yelp/mrjob
mrjob is a Python library for writing and running Hadoop Streaming MapReduce jobs, with support for Amazon EMR, Google Cloud Dataproc, self…
662613maintenance
apache/logging-flume
Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log-lik…
772567maintenance
justinzm/gopup
GoPUP is a Python library that provides convenient interfaces to a wide range of public Chinese data sources, including Baidu/Weibo/Google …
232548maintenance
geekyouth/SZT-bigdata
A big data passenger flow analysis system for the Shenzhen Metro, built on Shenzhen Tong smart-card swipe data. It demonstrates ETL and ana…
602475maintenance
cdarlint/winutils
A repository of precompiled Windows binaries (winutils.exe, hadoop.dll, hdfs.dll) needed to run Hadoop natively on Windows. It continues th…
322287maintenance
salesforce/TransmogrifAI
TransmogrifAI is an AutoML library written in Scala that runs on Apache Spark for building modular, reusable, strongly typed machine learni…
612277maintenance

← prev page 3 / 5 next →