Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: etl

671 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Pathway
Pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG applications. It provides a unified Py…
9762383active
pandas-dev/pandas
pandas is a fast, flexible Python library providing labeled data structures like the DataFrame for practical, real-world data analysis and …
9349564stable
apache/airflow
Apache Airflow is an open-source platform for programmatically authoring, scheduling, and monitoring workflows as directed acyclic graphs (…
9846613stable
apache/spark
Apache Spark is a unified analytics engine for large-scale data processing, providing high-level APIs in Scala, Java, Python, and R over an…
7743882stable
pola-rs/polars
Polars is an extremely fast analytical query engine for DataFrames written in Rust, with multi-threaded vectorized execution, lazy query op…
9539505stable
apache/kafka
Apache Kafka is an open-source distributed event streaming platform for building high-performance data pipelines, streaming analytics, and …
7733632stable
alibaba/canal
Canal is an Alibaba open-source component that parses MySQL binlog to provide incremental data subscription and consumption. It masquerades…
6629725stable
kestra-io/kestra
Kestra is an open-source, event-driven orchestration and scheduling platform for data, AI, and infrastructure workflows, written in Java. I…
9527933stable
apache/flink
Apache Flink is an open-source distributed stream processing framework for stateful computations over unbounded and bounded data streams, w…
7726294stable
vectordotdev/vector
Vector is a high-performance observability data pipeline written in Rust that collects, transforms, and routes logs, metrics, and traces to…
9522460active
airbytehq/airbyte
Airbyte is an open-source data movement platform providing 600+ connectors for replicating data from APIs, databases, and files into wareho…
7921960active
huggingface/datasets
Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr…
9821870stable
spotify/luigi
Luigi is a Python package for building complex pipelines of long-running batch jobs. It handles dependency resolution, workflow management,…
9118765stable
subquery/subql
SubQuery is an open-source, multi-chain data indexing framework for web3 that lets developers extract, transform, and query blockchain data…
7618753active
influxdata/telegraf
Telegraf is an open-source agent for collecting, processing, aggregating, and writing metrics, logs, and other telemetry data. It ships as …
9817767stable
alibaba/DataX
DataX is Alibaba's open-source offline data synchronization framework, the open version of Alibaba Cloud DataWorks data integration. It syn…
6417328stable
apache/arrow
Apache Arrow is a language-independent columnar in-memory data format specification plus multi-language libraries (C++, Python/PyArrow, Jav…
9417062stable
prestodb/presto
Presto is a distributed SQL query engine for running interactive queries against large datasets across heterogeneous data sources such as H…
9616724stable
dagster-io/dagster
Dagster is an open-source Python data orchestration platform for developing, producing, and observing data assets, with integrated lineage,…
9516067active
apache/hadoop
Apache Hadoop is an open-source framework for reliable, scalable distributed storage (HDFS) and processing (MapReduce, YARN) of large data …
7715640stable
Unstructured-IO/unstructured
Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into…
9515349active
elastic/logstash
Logstash is an open-source server-side data processing pipeline that ingests data from multiple sources simultaneously, transforms it, and …
9514925stable
ConardLi/easy-dataset
Easy Dataset is a self-hosted web application for building high-quality structured datasets for LLM fine-tuning, RAG, and model evaluation.…
7114833active
apache/dolphinscheduler
Apache DolphinScheduler is a modern data orchestration platform for building high-performance workflows with low-code drag-and-drop tooling…
9014447stable
Eclipse Deeplearning4J
Eclipse Deeplearning4J is an open-source deep learning framework and ecosystem for the JVM, including the ND4J linear algebra library, the …
7714246active
Dask
Dask is a flexible parallel and distributed computing library for Python that scales pandas, NumPy, scikit-learn, and other PyData tools to…
9913896stable
Fluentd
Fluentd is an open-source data collector that unifies log and event collection from many sources and routes them to files, databases, cloud…
9213579stable
Debezium
Debezium is an open source distributed platform for change data capture (CDC) that streams row-level database changes as events, most commo…
7713049stable
SpartnerNL/Laravel-Excel
Laravel Excel (maatwebsite/excel) is a Laravel package providing an elegant wrapper around PhpSpreadsheet for exporting and importing Excel…
9912705stable
elastic/beats
Elastic Beats is a family of lightweight data shippers written in Go that collect logs, metrics, network packets, audit data, and uptime si…
9512640active
datahub-project/datahub
DataHub is an open-source data catalog and metadata platform providing data discovery, governance, lineage, and observability across an org…
9812586active
OpenRefine/OpenRefine
OpenRefine is a free, open-source desktop application (running a local web server accessed via browser) for cleaning, transforming, and enr…
9011954stable
fivetran/great_expectations
Great Expectations (GX Core) is a Python library for validating, documenting, and profiling data using declarative, human-readable assertio…
9511740active
yutiansut/QUANTAXIS
QUANTAXIS is a Python quantitative finance framework providing local data fetching, backtesting, simulated and live trading, visualization,…
5411045active
kedro-org/kedro
Kedro is an open-source Python framework for building production-ready data engineering and data science pipelines using software engineeri…
9210971stable
getlago/lago
Lago is an open-source metering and usage-based billing platform that ingests usage events, manages subscriptions and pricing, and generate…
9510411active
modin-project/modin
Modin is a drop-in replacement for pandas that scales DataFrame operations across all CPU cores using Ray, Dask, or Unidist as execution en…
6710390active
EpistasisLab/tpot
TPOT (Tree-based Pipeline Optimization Tool) is a Python automated machine learning library that optimizes scikit-learn machine learning pi…
4410052active
johnkerl/miller
Miller (mlr) is a command-line tool like awk, sed, cut, join, and sort but designed for name-indexed data such as CSV, TSV, JSON, JSON Line…
9410003stable
NVIDIA/cudf
cuDF is a GPU-accelerated DataFrame library for tabular data processing, part of NVIDIA's RAPIDS suite. It provides a pandas-compatible Pyt…
959734stable
dataabc/weiboSpider
A Python command-line crawler that scrapes posts and profile data from Sina Weibo users, writing results to txt/csv/json files or MySQL/Mon…
629695active
apache/seatunnel
Apache SeaTunnel is a distributed, high-performance data integration platform for synchronizing massive amounts of data across hundreds of …
829587stable
dotnet/machinelearning
ML.NET is an open-source, cross-platform machine learning framework for .NET that lets developers build, train, and deploy custom ML models…
809351active
blue-yonder/tsfresh
tsfresh is a Python package that automatically extracts hundreds of features from time series using algorithms from statistics, signal proc…
809299stable
risingwavelabs/risingwave
RisingWave is an event streaming platform that continuously ingests data from databases, event streams, and webhooks, processes it incremen…
999290active
Delta Lake
Delta Lake is an open-source storage framework and table format that brings ACID transactions, scalable metadata handling, schema enforceme…
958960stable
mage-ai/mage-ai
Mage OSS is a self-hosted data pipeline tool for building, running, and orchestrating ETL/ELT workflows in a notebook-style UI using Python…
788814active
redpanda-data/connect
Redpanda Connect (formerly Benthos) is a declarative stream processor that moves data between hundreds of sources and sinks with transforma…
958736active
apache/beam
Apache Beam is an open-source unified programming model and SDK set (Java, Python, Go, SQL, TypeScript) for defining batch and streaming da…
938650stable
pentaho/pentaho-kettle
Pentaho Data Integration (Kettle) is an open-source ETL application for designing and running data extraction, transformation, and loading …
678382active
elasticsearch-dump/elasticsearch-dump
A Node.js CLI tool (elasticdump) for importing and exporting Elasticsearch and OpenSearch indices, moving data between clusters or to/from …
927939active
turbot/steampipe
Steampipe is a zero-ETL CLI tool that exposes APIs and cloud services as relational database tables so they can be queried live with standa…
987934active
OpenDCAI/DataFlow
DataFlow is a Python framework for preparing, cleaning, and synthesizing data for LLM training and RAG using composable LLM-based operators…
807764active
alteryx/featuretools
Featuretools is an open-source Python library for automated feature engineering, using Deep Feature Synthesis (DFS) to generate features fr…
657666active
enso-org/enso
Enso Analytics is a self-service data preparation and analysis platform for data teams, built on a hybrid visual/textual functional program…
917439active
microsoft/graphrag
GraphRAG is a Python library and CLI pipeline from Microsoft Research that extracts a knowledge graph from unstructured text, builds commun…
9135699maintenance
AlaSQL/alasql
AlaSQL is a lightweight in-memory SQL database written in pure JavaScript that runs in the browser, Node.js, and mobile apps. It queries bo…
997281active
feast-dev/feast
Feast is an open-source feature store for machine learning that manages offline stores for historical training data and low-latency online …
997230stable
snowplow/snowplow
Snowplow is a Customer Data Infrastructure platform that collects, validates, enriches, and streams event-level behavioral data from web, m…
637029stable
datajuicer/data-juicer
Data-Juicer is a Python library and data processing system for cleaning, deduplicating, synthesizing, and analyzing data for foundation mod…
946938active
apache/zeppelin
Apache Zeppelin is a web-based notebook for interactive, data-driven analytics and collaborative documents. It supports SQL, Scala, Python,…
776655stable
ibis-project/ibis
Ibis is a portable Python dataframe library that provides a single lazy dataframe API across more than 20 execution backends including Duck…
896643active
airweave-ai/airweave
Airweave is an open-source context retrieval layer that connects to apps, databases, and documents and exposes their data through a single …
746564active
dimitri/pgloader
pgloader is a data loading and migration tool for PostgreSQL built on the COPY command. It migrates entire databases from MySQL, SQLite, MS…
826504active
cloudquery/cloudquery
CloudQuery is an open-source CLI and plugin-based data pipeline tool that extracts cloud infrastructure configuration and security data fro…
956496active
nteract/papermill
Papermill is a Python library and CLI tool for parameterizing, executing, and analyzing Jupyter notebooks. It enables running notebooks pro…
736477active
apache/flink-cdc
Flink CDC is a distributed streaming data integration tool built on Apache Flink that captures change data from databases like MySQL and Po…
836468active
wireservice/csvkit
csvkit is a suite of Python-based command-line tools for converting to and working with CSV files. It includes utilities like csvcut, csvst…
756409stable
MaterializeInc/materialize
Materialize is a streaming database that uses incremental computation to maintain always-fresh, strongly consistent SQL views over data fro…
996361active
pachyderm/pachyderm
Pachyderm is a data-centric pipeline platform that automates data transformations with built-in data versioning and lineage tracking. It ru…
366308active
apache/hudi
Apache Hudi is an open data lakehouse platform built on a high-performance open table format that brings database functionality like transa…
916219stable
apache/nifi
Apache NiFi is an easy-to-use, powerful, and reliable system to process and distribute data, built on the JVM. It provides a browser-based …
946208stable
WeiYe-Jing/datax-web
DataX-Web is a distributed data synchronization tool built on top of Alibaba's DataX, providing a web UI to visually configure and manage d…
236018stable
apache/hive
Apache Hive is a distributed, fault-tolerant data warehouse system that enables reading, writing, and managing petabytes of data in distrib…
776015active
dlt-hub/dlt
dlt (data load tool) is an open-source Python library for building ELT data pipelines that extract data from REST APIs, SQL databases, clou…
985778stable
treeverse/lakeFS
lakeFS is an open-source data version control service that turns object storage (S3, Azure Blob, GCS) into a Git-like repository with branc…
945496stable
fluvio-community/fluvio
Fluvio is a distributed data streaming engine written in Rust that combines a Kafka-like event streaming platform with the Stateful DataFlo…
795247active
neo4j-labs/llm-graph-builder
A web application that transforms unstructured data (PDFs, DOCs, TXTs, YouTube videos, web pages) into a knowledge graph stored in Neo4j us…
905197active
tidyverse/dplyr
dplyr is an R package providing a consistent grammar of data manipulation with verbs like mutate(), select(), filter(), summarise(), and ar…
815062stable
javascriptdata/danfojs
Danfo.js is a JavaScript library providing high-performance, intuitive data structures (Series and DataFrame) for manipulating and processi…
575056active
jitsucom/jitsu
Jitsu is an open-source, self-hostable event data platform and Segment alternative that collects event data from websites, apps, and server…
975043active
ArroyoSystems/arroyo
Arroyo is a distributed stream processing engine written in Rust that lets users run stateful computations on high-volume real-time data st…
785016active
clidey/whodb
WhoDB is a lightweight, self-hosted database management workspace that lets teams explore schemas, edit data, run queries, and visualize re…
865014active
deepseek-ai/smallpond
Smallpond is a lightweight distributed data processing framework built on DuckDB and DeepSeek's 3FS shared file system. It lets users proce…
245000active
pudo/dataset
dataset is a Python library that makes reading and writing data in SQL databases as simple as working with JSON files. It provides implicit…
784871stable
amundsen-io/amundsen
Amundsen is an open-source data discovery and metadata engine that indexes data resources such as tables, dashboards, and streams, and powe…
674782active
feyninc/chonkie
Chonkie is a lightweight Python library for chunking text in RAG pipelines, offering multiple chunking algorithms, embeddings, and integrat…
844704active
SPLWare/esProc
esProc SPL is a JVM-based scripting language for structured data computation, usable as a standalone data analysis tool or embedded computi…
774684active
dataabc/weibo-crawler
A Python crawler for Sina Weibo that scrapes user profiles and posts, exporting data to CSV, JSON, MySQL, MongoDB, or SQLite, and optionall…
744625active
yihong0618/running_page
A self-hosted personal running homepage that syncs workout data from services like Strava, Garmin, Nike Run Club, and GPX files, then visua…
754504active
rudderlabs/rudder-server
RudderStack's open-source event streaming server, a privacy- and security-focused Segment alternative written in Go. It collects customer e…
954477active
rom1504/img2dataset
A Python tool that downloads large sets of image URLs and packages them into machine learning datasets, with resizing and caption support. …
564443active
unionai-oss/pandera
Pandera is a Python library for statistical data validation of dataframe-like objects, supporting pandas, polars, PySpark, and others. It l…
964442stable
tair-opensource/RedisShake
RedisShake is a Go-based tool for migrating and transforming data between Valkey/Redis instances, supporting standalone, master-slave, sent…
944427active
apache/streampark
Apache StreamPark is a streaming application development framework and one-stop cloud-native real-time computing platform for Apache Flink …
734328stable
briefercloud/briefer
Briefer is a self-hostable web application combining Notion-like notebooks and dashboards for data analysis, supporting Markdown, Python, S…
384319active
zendesk/maxwell
Maxwell's Daemon is a change data capture (CDC) application that reads MySQL binlogs and emits row-level changes as JSON to Kafka, Kinesis,…
964258active
pydata/xarray
Xarray is a Python library that adds labeled dimensions, coordinates, and attributes on top of NumPy-like N-dimensional arrays and datasets…
974190stable
dotnetcore/DotnetSpider
DotnetSpider is a .NET Standard web crawling and scraping framework that is lightweight, efficient, and cross-platform. It supports distrib…
574138active
aws/aws-sdk-pandas
AWS SDK for pandas (awswrangler) is a Python library that extends pandas with high-level APIs for reading and writing data across AWS servi…
944118active

page 1 / 7 next →