function: etl
671 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| justlovemaki/CloudFlare-AI-Insight-Daily A Cloudflare Workers-based content aggregation and generation platform that daily curates AI industry news, trending open-source projects, … | 63 | 1779 | active |
| lf-edge/ekuiper LF Edge eKuiper is a lightweight stream processing engine for IoT edge devices, offering SQL-based and graph-based rule engine for real-tim… | 98 | 1731 | active |
| bespokelabsai/curator Bespoke Curator is a Python library for bulk LLM inference and scalable synthetic data curation for post-training and structured data extra… | 79 | 1719 | active |
| jldbc/pybaseball pybaseball is a Python package that scrapes and retrieves current and historical baseball statistics from sources like MLB Statcast (Baseba… | 50 | 1719 | active |
| 4paradigm/OpenMLDB OpenMLDB is an open-source machine learning database that acts as a feature platform, computing consistent features for both offline traini… | 67 | 1710 | active |
| br-g/openf1 OpenF1 is a free, open-source API providing real-time and historical Formula 1 data, including lap timings, car telemetry, driver info, and… | 71 | 1692 | active |
| osm2pgsql-dev/osm2pgsql osm2pgsql is a command-line ETL tool that imports OpenStreetMap data in OSM, PBF, and O5M formats into a PostgreSQL/PostGIS database. It su… | 86 | 1680 | stable |
| Bruin Bruin is an end-to-end data platform whose CLI combines data ingestion (via its ingestr tool), SQL/Python/R transformations, and data quali… | 91 | 1678 | active |
| Multiwoven/multiwoven Multiwoven is an open-source Reverse ETL and data activation platform that syncs data from warehouses like Snowflake, BigQuery, Redshift, a… | 90 | 1671 | active |
| mortada/fredapi fredapi is a Python wrapper around the Federal Reserve Bank of St. Louis FRED and ALFRED web services for fetching economic data. It return… | 52 | 1654 | active |
| skrub-data/skrub skrub is a Python library that facilitates machine learning with dataframes, providing preprocessing, encoding, and wrangling tools for tab… | 93 | 1650 | active |
| sunlabuiuc/PyHealth PyHealth is an open-source Python toolkit for clinical deep learning, unifying healthcare datasets (EHRs, physiological signals, medical im… | 91 | 1650 | active |
| huggingface/aisheets Hugging Face AI Sheets is an open-source no-code web application for building, enriching, and transforming datasets using AI models. It can… | 58 | 1642 | active |
| ariacom/Seal-Report Seal Report & Task is an open-source (MIT) reporting and business intelligence framework written in C# for .NET, enabling users to build, p… | 86 | 1632 | active |
| meta-llama/synthetic-data-kit A Python CLI tool from Meta for generating high-quality synthetic datasets to fine-tune LLMs. It follows a four-stage pipeline (ingest, cre… | 42 | 1632 | active |
| karlicoss/HPI HPI (Human Programming Interface) is a Python package ('my') that unifies personal data from social networks, reading, health, location, me… | 80 | 1628 | active |
| snorkel-team/snorkel Snorkel is a Python library for programmatically building and managing training data using weak supervision, letting users write labeling f… | 62 | 6001 | maintenance |
| OpenDota OpenDota is an open-source Dota 2 data platform that provides a REST API and web UI for match and player statistics. It ingests data from V… | 76 | 1627 | active |
| Snowflake-Labs/pg_lake pg_lake is a set of PostgreSQL extensions that turn Postgres into a lakehouse engine for Apache Iceberg tables and raw data lake files in o… | 82 | 1625 | active |
| lit26/finvizfinance A Python library that scrapes and downloads financial data from the FinViz website, returning stock fundamentals, technicals, charts, news,… | 79 | 1623 | active |
| nerevu/riko riko is a pure Python stream processing library modeled after Yahoo! Pipes, combining reusable, configuration-driven modular pipes with syn… | 88 | 1606 | active |
| cuducos/minha-receita Minha Receita is an open-source project that consolidates Brazilian Federal Revenue (Receita Federal) CNPJ company data from scattered CSV … | 10 | 1597 | active |
| paradigmxyz/cryo cryo is a Rust-based CLI tool (also available as a Python package) that extracts Ethereum/EVM blockchain data into parquet, csv, , or Pytho… | 19 | 1580 | active |
| getdozer/dozer Dozer is a real-time data movement tool written in Rust that captures change data (CDC) from sources like Postgres, MySQL, Snowflake, and K… | 23 | 1579 | active |
| quixio/quix-streams Quix Streams is a pure Python framework for building real-time data pipelines and event-driven applications on Apache Kafka using a Streami… | 97 | 1568 | active |
| dimitri/pgcopydb pgcopydb is a command-line tool that automates copying a PostgreSQL database between two running servers, combining pg_dump/pg_restore for … | 81 | 1549 | active |
| ilyakatz/data-migrate A Ruby gem for Rails that lets you run data migrations alongside schema migrations. Data migrations live in db/data and are tracked in a da… | 81 | 1549 | active |
| mosaicml/streaming StreamingDataset is a Python library from MosaicML for fast, accurate streaming of training data from cloud object storage (S3, GCS, Azure,… | 70 | 1547 | active |
| combust/mleap MLeap is a serialization format (Bundle.ML) and portable execution engine for machine learning pipelines, implemented in Scala with Python … | 95 | 1543 | active |
| finos/legend Legend is an end-to-end data platform from FINOS (originated at Goldman Sachs) covering the full data lifecycle: modeling, curation, queryi… | 77 | 1543 | active |
| tensorflow/gnn TensorFlow GNN is a Python library for building Graph Neural Networks on TensorFlow, including a GraphTensor type for heterogeneous graphs,… | 66 | 1541 | active |
| hi-primus/optimus Optimus is a Python library for agile data preparation that provides a unified API over pandas, Dask, cuDF, Dask-cuDF, Vaex, and PySpark. I… | 23 | 1536 | active |
| FinanceData/FinanceDataReader FinanceDataReader is a Python library and CLI tool for reading financial data such as stock listings, stock prices, indexes, exchange rates… | 69 | 1535 | active |
| uwdata/arquero Arquero is a JavaScript library for query processing and transformation of array-backed, column-oriented data tables. It offers a fluent, d… | 45 | 1534 | stable |
| BemiHQ/BemiDB BemiDB is an open-source analytical data warehouse that combines built-in data source connectors (like Fivetran) with a Postgres-compatible… | 55 | 1531 | active |
| pyper-dev/pyper Pyper is a pure-Python library for building concurrent and parallel data pipelines using functional programming patterns. It unifies thread… | 28 | 1518 | active |
| lqzhgood/Shmily Shmily is a large umbrella project for exporting and archiving personal chat and communication records from QQ, WeChat, SMS, call logs, pho… | 55 | 1507 | active |
| pyjanitor-devs/pyjanitor pyjanitor is a Python library providing clean, readable APIs for data cleaning on pandas DataFrames, inspired by the R package janitor. It … | 99 | 1498 | active |
| apache/inlong Apache InLong is a one-stop, full-scenario integration framework for massive data, supporting data ingestion, synchronization, and subscrip… | 90 | 1497 | stable |
| Vincentqyw/cv-arxiv-daily An automated daily digest of computer vision and robotics arXiv papers (SLAM, SFM, visual localization, keypoint detection, image matching,… | 77 | 1494 | active |
| Surfer-Org/Protocol Surfer Protocol is an open-source framework for exporting personal data from platforms like Gmail, iMessages, Twitter, Notion, and ChatGPT.… | 13 | 1475 | active |
| cube2222/octosql OctoSQL is a Go CLI tool that lets you query, join, and transform data from multiple databases and file formats using SQL through a unified… | 23 | 5264 | maintenance |
| sfirke/janitor janitor is an R package with simple, user-friendly functions for examining and cleaning dirty data, such as formatting data.frame column na… | 23 | 1455 | stable |
| dadoonet/fscrawler FSCrawler is a Java-based file system crawler that indexes binary documents (PDF, MS Office, Open Office) into Elasticsearch, tracking new,… | 88 | 1450 | active |
| apache/hop Apache Hop is an open-source data and metadata orchestration platform for visually designing and running data integration pipelines and wor… | 94 | 1447 | active |
| tidyverse/tidyr tidyr is an R package from the tidyverse that provides tools for reshaping messy data into tidy form, including pivoting between long and w… | 70 | 1438 | stable |
| datazip-inc/olake OLake Go is a high-performance open-source EL (extract-load) engine written in Go that replicates databases (PostgreSQL, MySQL, MongoDB, Or… | 84 | 1431 | active |
| winedarksea/AutoTS AutoTS is a Python library for automated time series forecasting, offering dozens of sklearn-style models (statistical, ML, deep learning) … | 95 | 1430 | active |
| wgzhao/Addax Addax is a fast, extensible ETL tool for synchronizing data between heterogeneous SQL and NoSQL data sources, forked and evolved from Aliba… | 98 | 1425 | active |
| toluaina/pgsync PGSync is an open-source (MIT) change data capture tool that syncs PostgreSQL, MySQL, or MariaDB data to Elasticsearch or OpenSearch in rea… | 98 | 1415 | active |
| fmind/mlops-python-package A Python package template that provides a production-grade code base with MLOps best practices for building and deploying machine learning … | 91 | 1415 | active |
| amphi-ai/amphi-etl Amphi is a visual data preparation and ETL tool that lets users build data pipelines through a drag-and-drop interface while generating sta… | 70 | 1401 | active |
| data-forge/data-forge-ts Data-Forge is a TypeScript data transformation and analysis toolkit for JavaScript, inspired by Pandas and LINQ. It provides a DataFrame/Se… | 69 | 1392 | active |
| SebastienZh/StockTradebyZ A semi-automated stock screening application for China's A-share market that fetches daily K-line data via Tushare, applies quantitative pr… | 51 | 1390 | active |
| thinh-vu/vnstock Vnstock is an open-source Python library for extracting and analyzing Vietnam stock market data, returning data as pandas DataFrames via si… | 90 | 1383 | active |
| okfn-brasil/querido-diario Querido Diário is an open-source project by Open Knowledge Brasil that scrapes and aggregates Brazilian municipal official gazettes (diário… | 75 | 1373 | active |
| locationtech/geotrellis GeoTrellis is a Scala library and framework for high-performance reading, writing, and processing of geospatial raster and vector data. It … | 84 | 1372 | stable |
| spatie/simple-excel A PHP library for reading and writing simple Excel (xlsx) and CSV files with a fluent API. It uses generators and LazyCollections to keep m… | 86 | 1368 | active |
| nextstrain/ncov A Nextstrain build pipeline that analyzes SARS-CoV-2 viral genomes to understand their evolution and spread, producing phylogenetic visuali… | 88 | 1363 | active |
| nf-core/rnaseq nf-core/rnaseq is a Nextflow-based bioinformatics pipeline for analyzing RNA sequencing data, performing QC, trimming, (pseudo-)alignment w… | 94 | 1357 | active |
| SpiderClub/weibospider A distributed web crawler for Sina Weibo (Chinese microblogging platform) built with Python, Celery, and requests. It scrapes user profiles… | 32 | 4793 | maintenance |
| duckdb/dbt-duckdb dbt-duckdb is the dbt adapter plugin for DuckDB, an embedded OLAP database. It lets you run dbt data transformation pipelines in SQL or Pyt… | 94 | 1341 | active |
| datavane/tis TIS is an AI-native data integration platform built on DataX, Flink, and Flink-CDC that provides a visual Web-UI for zero-code batch and re… | 90 | 1339 | active |
| subsquid/squid-sdk Squid SDK is a TypeScript ETL toolkit for building blockchain data indexers, supporting Ethereum-like chains, Substrate, and Solana, with d… | 98 | 1337 | active |
| rwynn/monstache Monstache is a Go daemon that continuously syncs MongoDB collections into Elasticsearch (or OpenSearch) in realtime using change streams. I… | 48 | 1331 | active |
| petl-developers/petl petl is a general-purpose Python package for extracting, transforming and loading tables of data. It provides a lightweight, pure-Python to… | 98 | 1316 | stable |
| GoogleCloudPlatform/DataflowTemplates A collection of Google-provided Apache Beam pipeline templates for Google Cloud Dataflow that solve common in-cloud data tasks like import/… | 95 | 1310 | active |
| Mojang/DataFixerUpper DataFixerUpper is a Java library from Mojang for incrementally building, merging, and optimizing data transformations between schema versio… | 59 | 1309 | active |
| arkflow-rs/arkflow ArkFlow is a high-performance stream processing engine written in Rust on top of Tokio, connecting configurable inputs (Kafka, MQTT, HTTP, … | 70 | 1302 | active |
| elixir-explorer/explorer Explorer is an Elixir library providing series (one-dimensional) and dataframes (two-dimensional) for fast data exploration and manipulatio… | 87 | 1291 | active |
| DTStack/Taier Taier is a self-hosted distributed dispatching platform for big data that handles task submission, DAG-based scheduling, and operations/mai… | 84 | 1284 | active |
| Nixtla/mlforecast mlforecast is a Python framework for time series forecasting using any machine learning model with fit/predict methods, providing efficient… | 96 | 1269 | active |
| meta-pytorch/data TorchData is a PyTorch library providing scalable, performant data loading utilities, including StatefulDataLoader, a drop-in replacement f… | 67 | 1263 | active |
| slothflowlabs/duckle Duckle is an open-source ETL/ELT platform built on DuckDB that you self-host on your own servers or cloud, with a visual no-code/low-code c… | 81 | 1262 | active |
| apache/datafusion-comet Apache DataFusion Comet is a high-performance accelerator plugin for Apache Spark that keeps queries Arrow-native end-to-end, executing ope… | 81 | 1262 | active |
| 0xSero/ai-data-extraction A Python toolkit of extraction scripts that pulls complete chat, agent, and code-context history from AI coding assistants like Cursor, Cla… | 60 | 1257 | active |
| firecrawl/fire-enrich Fire Enrich is an AI-powered data enrichment web application that transforms a list of email addresses into rich company datasets, includin… | 39 | 1255 | active |
| astronomer/astronomer-cosmos Astronomer Cosmos is an open-source Python library that converts dbt Core or dbt Fusion projects into Apache Airflow DAGs and task groups w… | 98 | 1254 | active |
| sentinel-hub/eo-learn eo-learn is a collection of open-source Python packages for accessing and processing spatio-temporal satellite imagery, built around modula… | 51 | 1247 | active |
| zinggAI/zingg Zingg is an ML-based tool for scalable master data management, entity resolution, identity resolution, and record deduplication. It runs on… | 88 | 1243 | active |
| darold/ora2pg Ora2Pg is a free Perl-based tool that migrates an Oracle database to PostgreSQL by scanning the source database and generating SQL scripts … | 66 | 1231 | active |
| thetahealth/mirobody Mirobody is an open-source, AI-native health data engine that collects readings from lab reports, wearables, and genomics, standardizes the… | 84 | 1229 | active |
| kevwan/go-stash go-stash is a high-performance, open-source server-side data processing pipeline written in Go that ingests data from Kafka, applies config… | 60 | 1224 | active |
| marcboeker/gmail-to-sqlite A Python CLI application that syncs Gmail messages into a local SQLite database, supporting incremental and full syncs with deletion detect… | 56 | 1221 | active |
| skyplane-project/skyplane Skyplane is a CLI tool for blazingly fast bulk data transfers between cloud object stores (AWS S3, Azure Blob, GCS, IBM COS) and local disk… | 23 | 1216 | active |
| pytroll/satpy Satpy is a Python library for reading, manipulating, and writing meteorological remote sensing data from earth-observing satellites. It sup… | 86 | 1205 | active |
| apache/incubator-xtable Apache XTable (incubating) is a cross-table converter that translates lakehouse table format metadata between Apache Hudi, Apache Iceberg, … | 84 | 1205 | active |
| gityuanbao/share A personal open-source repository whose main component is akshare_collector, a Python tool built on AKShare that collects Chinese financial… | 64 | 1201 | active |
| xorbitsai/xorbits Xorbits is an open-source distributed computing framework that scales Python data science and machine learning workloads from a laptop to l… | 64 | 1199 | active |
| wireservice/agate agate is a Python data analysis library optimized for humans instead of machines, offering a readable alternative to numpy and pandas. It p… | 75 | 1198 | stable |
| marsupialtail/quokka Quokka is a lightweight distributed dataflow/query engine written in Python, built on Ray, DuckDB, Polars, and Arrow, designed for stateful… | 23 | 1192 | active |
| nextgenhealthcare/connect Mirth Connect by NextGen Healthcare is an open-source healthcare integration engine that filters, transforms, extracts, and routes messages… | 23 | 1190 | active |
| go-mysql-org/go-mysql-elasticsearch A Go service that automatically syncs MySQL data into Elasticsearch, using mysqldump for initial load and binlog parsing for incremental up… | 32 | 4148 | maintenance |
| xataio/pgstream pgstream is an open-source change data capture (CDC) tool and Go library that replicates PostgreSQL data, including DDL/schema changes, to … | 88 | 1186 | active |
| functime-org/functime functime is a Python library for production-ready global forecasting and time-series feature extraction on large panel datasets, built on l… | 70 | 1183 | active |
| machow/siuba siuba is a Python library that ports R's dplyr syntax to pandas DataFrames and SQL databases, using pipe (>>) chaining and lazy expressions… | 42 | 1183 | active |
| apache/amoro Apache Amoro (incubating) is a Lakehouse management system built on open data lake formats like Iceberg, Paimon, and Mixed-Hive. It provide… | 73 | 1171 | active |
| lhotse-speech/lhotse Lhotse is a Python library for flexible, scalable preparation of multimodal (speech, audio, video, image, text) data for machine learning, … | 89 | 1149 | active |
| Tavish9/any4lerobot Any4LeRobot is a curated collection of Python utilities for the Hugging Face LeRobot robotics ecosystem, including dataset conversion scrip… | 64 | 1140 | active |
| certtools/intelmq IntelMQ is an open-source solution for IT security teams (CERTs, CSIRTs, SOCs) for collecting and processing security feeds using a message… | 63 | 1133 | active |