function: etl
671 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Pathway Pathway is a Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG applications. It provides a unified Py… | 97 | 62383 | active |
| pandas-dev/pandas pandas is a fast, flexible Python library providing labeled data structures like the DataFrame for practical, real-world data analysis and … | 93 | 49564 | stable |
| apache/airflow Apache Airflow is an open-source platform for programmatically authoring, scheduling, and monitoring workflows as directed acyclic graphs (… | 98 | 46613 | stable |
| apache/spark Apache Spark is a unified analytics engine for large-scale data processing, providing high-level APIs in Scala, Java, Python, and R over an… | 77 | 43882 | stable |
| pola-rs/polars Polars is an extremely fast analytical query engine for DataFrames written in Rust, with multi-threaded vectorized execution, lazy query op… | 95 | 39505 | stable |
| apache/kafka Apache Kafka is an open-source distributed event streaming platform for building high-performance data pipelines, streaming analytics, and … | 77 | 33632 | stable |
| alibaba/canal Canal is an Alibaba open-source component that parses MySQL binlog to provide incremental data subscription and consumption. It masquerades… | 66 | 29725 | stable |
| kestra-io/kestra Kestra is an open-source, event-driven orchestration and scheduling platform for data, AI, and infrastructure workflows, written in Java. I… | 95 | 27933 | stable |
| apache/flink Apache Flink is an open-source distributed stream processing framework for stateful computations over unbounded and bounded data streams, w… | 77 | 26294 | stable |
| vectordotdev/vector Vector is a high-performance observability data pipeline written in Rust that collects, transforms, and routes logs, metrics, and traces to… | 95 | 22460 | active |
| airbytehq/airbyte Airbyte is an open-source data movement platform providing 600+ connectors for replicating data from APIs, databases, and files into wareho… | 79 | 21960 | active |
| huggingface/datasets Hugging Face Datasets is a Python library providing one-line access to hundreds of thousands of public datasets on the Hugging Face Hub acr… | 98 | 21870 | stable |
| spotify/luigi Luigi is a Python package for building complex pipelines of long-running batch jobs. It handles dependency resolution, workflow management,… | 91 | 18765 | stable |
| subquery/subql SubQuery is an open-source, multi-chain data indexing framework for web3 that lets developers extract, transform, and query blockchain data… | 76 | 18753 | active |
| influxdata/telegraf Telegraf is an open-source agent for collecting, processing, aggregating, and writing metrics, logs, and other telemetry data. It ships as … | 98 | 17767 | stable |
| alibaba/DataX DataX is Alibaba's open-source offline data synchronization framework, the open version of Alibaba Cloud DataWorks data integration. It syn… | 64 | 17328 | stable |
| apache/arrow Apache Arrow is a language-independent columnar in-memory data format specification plus multi-language libraries (C++, Python/PyArrow, Jav… | 94 | 17062 | stable |
| prestodb/presto Presto is a distributed SQL query engine for running interactive queries against large datasets across heterogeneous data sources such as H… | 96 | 16724 | stable |
| dagster-io/dagster Dagster is an open-source Python data orchestration platform for developing, producing, and observing data assets, with integrated lineage,… | 95 | 16067 | active |
| apache/hadoop Apache Hadoop is an open-source framework for reliable, scalable distributed storage (HDFS) and processing (MapReduce, YARN) of large data … | 77 | 15640 | stable |
| Unstructured-IO/unstructured Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into… | 95 | 15349 | active |
| elastic/logstash Logstash is an open-source server-side data processing pipeline that ingests data from multiple sources simultaneously, transforms it, and … | 95 | 14925 | stable |
| ConardLi/easy-dataset Easy Dataset is a self-hosted web application for building high-quality structured datasets for LLM fine-tuning, RAG, and model evaluation.… | 71 | 14833 | active |
| apache/dolphinscheduler Apache DolphinScheduler is a modern data orchestration platform for building high-performance workflows with low-code drag-and-drop tooling… | 90 | 14447 | stable |
| Eclipse Deeplearning4J Eclipse Deeplearning4J is an open-source deep learning framework and ecosystem for the JVM, including the ND4J linear algebra library, the … | 77 | 14246 | active |
| Dask Dask is a flexible parallel and distributed computing library for Python that scales pandas, NumPy, scikit-learn, and other PyData tools to… | 99 | 13896 | stable |
| Fluentd Fluentd is an open-source data collector that unifies log and event collection from many sources and routes them to files, databases, cloud… | 92 | 13579 | stable |
| Debezium Debezium is an open source distributed platform for change data capture (CDC) that streams row-level database changes as events, most commo… | 77 | 13049 | stable |
| SpartnerNL/Laravel-Excel Laravel Excel (maatwebsite/excel) is a Laravel package providing an elegant wrapper around PhpSpreadsheet for exporting and importing Excel… | 99 | 12705 | stable |
| elastic/beats Elastic Beats is a family of lightweight data shippers written in Go that collect logs, metrics, network packets, audit data, and uptime si… | 95 | 12640 | active |
| datahub-project/datahub DataHub is an open-source data catalog and metadata platform providing data discovery, governance, lineage, and observability across an org… | 98 | 12586 | active |
| OpenRefine/OpenRefine OpenRefine is a free, open-source desktop application (running a local web server accessed via browser) for cleaning, transforming, and enr… | 90 | 11954 | stable |
| fivetran/great_expectations Great Expectations (GX Core) is a Python library for validating, documenting, and profiling data using declarative, human-readable assertio… | 95 | 11740 | active |
| yutiansut/QUANTAXIS QUANTAXIS is a Python quantitative finance framework providing local data fetching, backtesting, simulated and live trading, visualization,… | 54 | 11045 | active |
| kedro-org/kedro Kedro is an open-source Python framework for building production-ready data engineering and data science pipelines using software engineeri… | 92 | 10971 | stable |
| getlago/lago Lago is an open-source metering and usage-based billing platform that ingests usage events, manages subscriptions and pricing, and generate… | 95 | 10411 | active |
| modin-project/modin Modin is a drop-in replacement for pandas that scales DataFrame operations across all CPU cores using Ray, Dask, or Unidist as execution en… | 67 | 10390 | active |
| EpistasisLab/tpot TPOT (Tree-based Pipeline Optimization Tool) is a Python automated machine learning library that optimizes scikit-learn machine learning pi… | 44 | 10052 | active |
| johnkerl/miller Miller (mlr) is a command-line tool like awk, sed, cut, join, and sort but designed for name-indexed data such as CSV, TSV, JSON, JSON Line… | 94 | 10003 | stable |
| NVIDIA/cudf cuDF is a GPU-accelerated DataFrame library for tabular data processing, part of NVIDIA's RAPIDS suite. It provides a pandas-compatible Pyt… | 95 | 9734 | stable |
| dataabc/weiboSpider A Python command-line crawler that scrapes posts and profile data from Sina Weibo users, writing results to txt/csv/json files or MySQL/Mon… | 62 | 9695 | active |
| apache/seatunnel Apache SeaTunnel is a distributed, high-performance data integration platform for synchronizing massive amounts of data across hundreds of … | 82 | 9587 | stable |
| dotnet/machinelearning ML.NET is an open-source, cross-platform machine learning framework for .NET that lets developers build, train, and deploy custom ML models… | 80 | 9351 | active |
| blue-yonder/tsfresh tsfresh is a Python package that automatically extracts hundreds of features from time series using algorithms from statistics, signal proc… | 80 | 9299 | stable |
| risingwavelabs/risingwave RisingWave is an event streaming platform that continuously ingests data from databases, event streams, and webhooks, processes it incremen… | 99 | 9290 | active |
| Delta Lake Delta Lake is an open-source storage framework and table format that brings ACID transactions, scalable metadata handling, schema enforceme… | 95 | 8960 | stable |
| mage-ai/mage-ai Mage OSS is a self-hosted data pipeline tool for building, running, and orchestrating ETL/ELT workflows in a notebook-style UI using Python… | 78 | 8814 | active |
| redpanda-data/connect Redpanda Connect (formerly Benthos) is a declarative stream processor that moves data between hundreds of sources and sinks with transforma… | 95 | 8736 | active |
| apache/beam Apache Beam is an open-source unified programming model and SDK set (Java, Python, Go, SQL, TypeScript) for defining batch and streaming da… | 93 | 8650 | stable |
| pentaho/pentaho-kettle Pentaho Data Integration (Kettle) is an open-source ETL application for designing and running data extraction, transformation, and loading … | 67 | 8382 | active |
| elasticsearch-dump/elasticsearch-dump A Node.js CLI tool (elasticdump) for importing and exporting Elasticsearch and OpenSearch indices, moving data between clusters or to/from … | 92 | 7939 | active |
| turbot/steampipe Steampipe is a zero-ETL CLI tool that exposes APIs and cloud services as relational database tables so they can be queried live with standa… | 98 | 7934 | active |
| OpenDCAI/DataFlow DataFlow is a Python framework for preparing, cleaning, and synthesizing data for LLM training and RAG using composable LLM-based operators… | 80 | 7764 | active |
| alteryx/featuretools Featuretools is an open-source Python library for automated feature engineering, using Deep Feature Synthesis (DFS) to generate features fr… | 65 | 7666 | active |
| enso-org/enso Enso Analytics is a self-service data preparation and analysis platform for data teams, built on a hybrid visual/textual functional program… | 91 | 7439 | active |
| microsoft/graphrag GraphRAG is a Python library and CLI pipeline from Microsoft Research that extracts a knowledge graph from unstructured text, builds commun… | 91 | 35699 | maintenance |
| AlaSQL/alasql AlaSQL is a lightweight in-memory SQL database written in pure JavaScript that runs in the browser, Node.js, and mobile apps. It queries bo… | 99 | 7281 | active |
| feast-dev/feast Feast is an open-source feature store for machine learning that manages offline stores for historical training data and low-latency online … | 99 | 7230 | stable |
| snowplow/snowplow Snowplow is a Customer Data Infrastructure platform that collects, validates, enriches, and streams event-level behavioral data from web, m… | 63 | 7029 | stable |
| datajuicer/data-juicer Data-Juicer is a Python library and data processing system for cleaning, deduplicating, synthesizing, and analyzing data for foundation mod… | 94 | 6938 | active |
| apache/zeppelin Apache Zeppelin is a web-based notebook for interactive, data-driven analytics and collaborative documents. It supports SQL, Scala, Python,… | 77 | 6655 | stable |
| ibis-project/ibis Ibis is a portable Python dataframe library that provides a single lazy dataframe API across more than 20 execution backends including Duck… | 89 | 6643 | active |
| airweave-ai/airweave Airweave is an open-source context retrieval layer that connects to apps, databases, and documents and exposes their data through a single … | 74 | 6564 | active |
| dimitri/pgloader pgloader is a data loading and migration tool for PostgreSQL built on the COPY command. It migrates entire databases from MySQL, SQLite, MS… | 82 | 6504 | active |
| cloudquery/cloudquery CloudQuery is an open-source CLI and plugin-based data pipeline tool that extracts cloud infrastructure configuration and security data fro… | 95 | 6496 | active |
| nteract/papermill Papermill is a Python library and CLI tool for parameterizing, executing, and analyzing Jupyter notebooks. It enables running notebooks pro… | 73 | 6477 | active |
| apache/flink-cdc Flink CDC is a distributed streaming data integration tool built on Apache Flink that captures change data from databases like MySQL and Po… | 83 | 6468 | active |
| wireservice/csvkit csvkit is a suite of Python-based command-line tools for converting to and working with CSV files. It includes utilities like csvcut, csvst… | 75 | 6409 | stable |
| MaterializeInc/materialize Materialize is a streaming database that uses incremental computation to maintain always-fresh, strongly consistent SQL views over data fro… | 99 | 6361 | active |
| pachyderm/pachyderm Pachyderm is a data-centric pipeline platform that automates data transformations with built-in data versioning and lineage tracking. It ru… | 36 | 6308 | active |
| apache/hudi Apache Hudi is an open data lakehouse platform built on a high-performance open table format that brings database functionality like transa… | 91 | 6219 | stable |
| apache/nifi Apache NiFi is an easy-to-use, powerful, and reliable system to process and distribute data, built on the JVM. It provides a browser-based … | 94 | 6208 | stable |
| WeiYe-Jing/datax-web DataX-Web is a distributed data synchronization tool built on top of Alibaba's DataX, providing a web UI to visually configure and manage d… | 23 | 6018 | stable |
| apache/hive Apache Hive is a distributed, fault-tolerant data warehouse system that enables reading, writing, and managing petabytes of data in distrib… | 77 | 6015 | active |
| dlt-hub/dlt dlt (data load tool) is an open-source Python library for building ELT data pipelines that extract data from REST APIs, SQL databases, clou… | 98 | 5778 | stable |
| treeverse/lakeFS lakeFS is an open-source data version control service that turns object storage (S3, Azure Blob, GCS) into a Git-like repository with branc… | 94 | 5496 | stable |
| fluvio-community/fluvio Fluvio is a distributed data streaming engine written in Rust that combines a Kafka-like event streaming platform with the Stateful DataFlo… | 79 | 5247 | active |
| neo4j-labs/llm-graph-builder A web application that transforms unstructured data (PDFs, DOCs, TXTs, YouTube videos, web pages) into a knowledge graph stored in Neo4j us… | 90 | 5197 | active |
| tidyverse/dplyr dplyr is an R package providing a consistent grammar of data manipulation with verbs like mutate(), select(), filter(), summarise(), and ar… | 81 | 5062 | stable |
| javascriptdata/danfojs Danfo.js is a JavaScript library providing high-performance, intuitive data structures (Series and DataFrame) for manipulating and processi… | 57 | 5056 | active |
| jitsucom/jitsu Jitsu is an open-source, self-hostable event data platform and Segment alternative that collects event data from websites, apps, and server… | 97 | 5043 | active |
| ArroyoSystems/arroyo Arroyo is a distributed stream processing engine written in Rust that lets users run stateful computations on high-volume real-time data st… | 78 | 5016 | active |
| clidey/whodb WhoDB is a lightweight, self-hosted database management workspace that lets teams explore schemas, edit data, run queries, and visualize re… | 86 | 5014 | active |
| deepseek-ai/smallpond Smallpond is a lightweight distributed data processing framework built on DuckDB and DeepSeek's 3FS shared file system. It lets users proce… | 24 | 5000 | active |
| pudo/dataset dataset is a Python library that makes reading and writing data in SQL databases as simple as working with JSON files. It provides implicit… | 78 | 4871 | stable |
| amundsen-io/amundsen Amundsen is an open-source data discovery and metadata engine that indexes data resources such as tables, dashboards, and streams, and powe… | 67 | 4782 | active |
| feyninc/chonkie Chonkie is a lightweight Python library for chunking text in RAG pipelines, offering multiple chunking algorithms, embeddings, and integrat… | 84 | 4704 | active |
| SPLWare/esProc esProc SPL is a JVM-based scripting language for structured data computation, usable as a standalone data analysis tool or embedded computi… | 77 | 4684 | active |
| dataabc/weibo-crawler A Python crawler for Sina Weibo that scrapes user profiles and posts, exporting data to CSV, JSON, MySQL, MongoDB, or SQLite, and optionall… | 74 | 4625 | active |
| yihong0618/running_page A self-hosted personal running homepage that syncs workout data from services like Strava, Garmin, Nike Run Club, and GPX files, then visua… | 75 | 4504 | active |
| rudderlabs/rudder-server RudderStack's open-source event streaming server, a privacy- and security-focused Segment alternative written in Go. It collects customer e… | 95 | 4477 | active |
| rom1504/img2dataset A Python tool that downloads large sets of image URLs and packages them into machine learning datasets, with resizing and caption support. … | 56 | 4443 | active |
| unionai-oss/pandera Pandera is a Python library for statistical data validation of dataframe-like objects, supporting pandas, polars, PySpark, and others. It l… | 96 | 4442 | stable |
| tair-opensource/RedisShake RedisShake is a Go-based tool for migrating and transforming data between Valkey/Redis instances, supporting standalone, master-slave, sent… | 94 | 4427 | active |
| apache/streampark Apache StreamPark is a streaming application development framework and one-stop cloud-native real-time computing platform for Apache Flink … | 73 | 4328 | stable |
| briefercloud/briefer Briefer is a self-hostable web application combining Notion-like notebooks and dashboards for data analysis, supporting Markdown, Python, S… | 38 | 4319 | active |
| zendesk/maxwell Maxwell's Daemon is a change data capture (CDC) application that reads MySQL binlogs and emits row-level changes as JSON to Kafka, Kinesis,… | 96 | 4258 | active |
| pydata/xarray Xarray is a Python library that adds labeled dimensions, coordinates, and attributes on top of NumPy-like N-dimensional arrays and datasets… | 97 | 4190 | stable |
| dotnetcore/DotnetSpider DotnetSpider is a .NET Standard web crawling and scraping framework that is lightweight, efficient, and cross-platform. It supports distrib… | 57 | 4138 | active |
| aws/aws-sdk-pandas AWS SDK for pandas (awswrangler) is a Python library that extends pandas with high-level APIs for reading and writing data across AWS servi… | 94 | 4118 | active |
page 1 / 7 next →