function: etl
671 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| nghuyong/WeiboSpider A continuously maintained Python web scraping tool for Sina Weibo built on Scrapy and the new weibo.com API. It collects user profiles, pos… | 73 | 4109 | active |
| DTStack/chunjun ChunJun (formerly FlinkX) is a distributed, batch-and-stream data integration framework built on Apache Flink. It synchronizes and computes… | 54 | 4099 | active |
| mahdibland/V2RayAggregator A Python-based automation that aggregates free proxy nodes (Shadowsocks, SSR, Trojan, Vmess) from public sources, deduplicates them, and sp… | 67 | 4008 | active |
| WikiExtractor/wikiextractor WikiExtractor is a Python command-line tool that extracts and cleans plain text from Wikipedia database backup dumps (XML). It supports tem… | 98 | 4001 | stable |
| ucbepic/docetl DocETL is a Python library and CLI for building LLM-powered data processing and ETL pipelines over structured and unstructured data using d… | 86 | 3995 | active |
| Rdatatable/data.table data.table is an R package providing a high-performance, memory-efficient replacement for base R's data.frame with a concise indexing synta… | 77 | 3913 | stable |
| Netflix/maestro Netflix Maestro is a general-purpose workflow orchestrator providing workflow-as-a-service for data, ML, and software pipelines. It schedul… | 69 | 3829 | active |
| dathere/qsv qsv is a blazing-fast, Rust-based command-line data-wrangling toolkit with 50+ composable commands for querying, transforming, validating, … | 98 | 3767 | active |
| jtablesaw/tablesaw Tablesaw is a Java dataframe library for loading, cleaning, transforming, filtering, and summarizing tabular data, with descriptive statist… | 86 | 3761 | active |
| awslabs/deequ Deequ is a Scala library built on Apache Spark for defining 'unit tests for data' that measure data quality in large datasets. It computes … | 93 | 3643 | active |
| ploomber/ploomber Ploomber is a Python framework for building maintainable data pipelines from scripts and Jupyter notebooks, with iterative local developmen… | 10 | 3622 | active |
| dtinit/data-transfer-project An open-source framework providing common data models and adapter libraries that enable direct service-to-service transfer of user data bet… | 67 | 3619 | active |
| waditu/tushare TuShare is a Python library for crawling, cleaning, and storing historical and realtime financial data for China stocks and futures. It pro… | 23 | 15366 | maintenance |
| chrislusf/gleam Gleam is a fast, efficient distributed map/reduce execution system written in pure Go, defining computations as DAG flows that can run stan… | 75 | 3562 | active |
| noflo/noflo NoFlo is a JavaScript implementation of Flow-Based Programming (FBP), where applications are defined as graphs of independent components co… | 64 | 3552 | active |
| nextflow-io/nextflow Nextflow is an open-source DSL and runtime for building scalable, portable, and reproducible data-driven computational pipelines, based on … | 97 | 3475 | stable |
| ankane/pgsync pgsync is a Ruby command-line tool that syncs data from one Postgres database to another, similar to pg_dump/pg_restore. It transfers table… | 76 | 3468 | active |
| arpanghosh8453/garmin-grafana A Dockerized Python application that fetches health and fitness data from Garmin Connect servers and stores it in a local InfluxDB time-ser… | 79 | 3432 | active |
| xlwings/xlwings xlwings is a BSD-licensed Python library that lets you call Python from Excel and automate Excel from Python, replacing VBA macros with Pyt… | 99 | 3396 | active |
| apache/paimon Apache Paimon is a lakehouse table format that combines lake format with LSM tree structure to support real-time streaming updates. It enab… | 87 | 3384 | active |
| linuxfoundation/crowd.dev LFX Community Data Platform (formerly crowd.dev) is a self-hostable platform that unifies community and developer engagement data from many… | 67 | 3364 | active |
| hicccc77/WeFlow WeFlow is a fully local desktop application for viewing, analyzing, and exporting WeChat chat history, including yearly data reports and vi… | 59 | 13959 | maintenance |
| django-import-export/django-import-export A Django application and library for importing and exporting data between Django models and file formats like CSV, XLSX, JSON, YAML, and HT… | 93 | 3334 | stable |
| lakehq/sail Sail is an open-source, Rust-native multimodal compute engine that serves as a drop-in replacement for Apache Spark, unifying batch process… | 93 | 3333 | active |
| huggingface/datatrove DataTrove is a Python library from Hugging Face for processing, filtering, and deduplicating large-scale text data through customizable pip… | 86 | 3308 | active |
| tcgoetz/GarminDB A Python package and CLI that downloads and parses health and activity data from Garmin Connect, Garmin watches, FitBit CSV, and MS Health … | 95 | 3277 | active |
| WeBankFinTech/DataSphereStudio DataSphere Studio is a one-stop data application development and management portal from WeBank, built on the Linkis computing middleware. I… | 50 | 3262 | active |
| SQLMesh/sqlmesh SQLMesh is a scalable data transformation framework for running and deploying SQL or Python data models, backward compatible with dbt. It p… | 93 | 3256 | active |
| PeerDB-io/peerdb PeerDB is an open-source ETL/CDC tool that streams data from Postgres to data warehouses, queues, and storage engines, with an actively mai… | 93 | 3252 | active |
| lakesoul-io/LakeSoul LakeSoul is a cloud-native, real-time lakehouse framework with a Rust-native core providing ACID table format, concurrent upserts, incremen… | 82 | 3247 | active |
| pydata/pandas-datareader A Python library that extracts data from remote internet sources such as FRED, Fama/French, World Bank, OECD, and Eurostat directly into pa… | 93 | 3238 | active |
| oxylabs/ai-crawler-py A Python client library for Oxylabs AI Studio's AI-Crawler, a commercial service that crawls websites from a starting URL, uses natural lan… | 61 | 3218 | active |
| duckdb/pg_duckdb pg_duckdb is an official PostgreSQL extension that embeds DuckDB's columnar-vectorized analytics engine inside Postgres. It lets users run … | 72 | 3210 | active |
| Wisser/Jailer Jailer is a Java-based tool for database subsetting and relational data browsing. It extracts small, referentially intact slices of product… | 99 | 3194 | active |
| webdataset/webdataset A Python library providing a high-performance sequential I/O system based on tar-shard files for large-scale deep learning training, with s… | 52 | 3169 | stable |
| omkarcloud/google-maps-scraper A desktop application and API for scraping Google Maps business data, extracting 50+ data points including emails, phone numbers, social pr… | 72 | 3130 | active |
| blockchain-etl/ethereum-etl A Python CLI tool that extracts Ethereum blockchain data (blocks, transactions, token transfers, logs, traces, contracts) and exports it to… | 52 | 3127 | active |
| apache/devlake Apache DevLake is an open-source dev data platform that ingests, analyzes, and visualizes fragmented data from DevOps tools like GitHub, Gi… | 93 | 3111 | active |
| alldatacenter/alldata AllData is a definable data platform (数据中台) that integrates open-source components like DolphinScheduler, DataX, SeaTunnel, OpenMetadata, a… | 83 | 3082 | active |
| TimelyDataflow/differential-dataflow A Rust library implementing differential dataflow on top of timely dataflow, enabling data-parallel programs that efficiently process large… | 93 | 2998 | active |
| spring-projects/spring-batch Spring Batch is a lightweight, comprehensive batch framework for Java and Spring that enables development of robust batch applications. It … | 95 | 2951 | stable |
| duckdb/ducklake DuckLake is an integrated data lake and lakehouse catalog format that stores data as Parquet files and manages metadata in an ACID-complian… | 65 | 2939 | active |
| cdhigh/KindleEar KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep… | 73 | 2865 | active |
| snakemake/snakemake Snakemake is a Python-based workflow management system for creating reproducible and scalable data analyses. Workflows are defined via a re… | 95 | 2853 | stable |
| numaproj/numaflow Numaflow is a Kubernetes-native, serverless platform for running massively parallel data and streaming processing jobs. It lets developers … | 98 | 2825 | active |
| datachain-ai/datachain DataChain is a Python library that turns unstructured files in S3, GCS, and Azure into typed, versioned datasets queryable at warehouse spe… | 86 | 2811 | active |
| dfinke/ImportExcel A PowerShell module for importing and exporting Excel spreadsheets without requiring Excel to be installed. It supports creating tables, pi… | 67 | 2733 | active |
| Ontos-AI/knowhere Knowhere is a document ingestion and parsing platform that transforms unstructured documents (PDF, DOCX, XLSX, PPTX, images) into structure… | 81 | 2673 | active |
| sfu-db/connector-x ConnectorX is a Rust-based library with Python bindings that loads data from databases into DataFrames (pandas, Arrow, Polars) as fast and … | 83 | 2646 | active |
| LLMQuant/quant-mind QuantMind is an agent-native Python framework that refines raw financial information such as papers, news, and SEC filings into typed, time… | 64 | 2630 | active |
| spotify/scio Scio is a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. It provides a unified batch and strea… | 96 | 2628 | active |
| OpenLineage/OpenLineage OpenLineage is an open standard and framework for collecting data lineage metadata, defining a generic model of jobs, runs, and datasets wi… | 97 | 2626 | active |
| dgunning/edgartools EdgarTools is a Python library for accessing and parsing SEC EDGAR filings as structured, typed Python objects. It provides a clean API for… | 94 | 2624 | active |
| meltano/meltano Meltano is an open-source, declarative, code-first data integration engine and CLI for building and running ELT pipelines. It orchestrates … | 97 | 2610 | stable |
| apache/hamilton Apache Hamilton is a lightweight Python library for defining directed acyclic graphs (DAGs) of data transformations as plain, testable Pyth… | 91 | 2575 | active |
| neilotoole/sq sq is a command-line data wrangling tool that provides jq-style querying over SQL databases and document formats like CSV, Excel, and JSON.… | 96 | 2555 | active |
| graphistry/pygraphistry PyGraphistry is a Python library for loading, shaping, and visually exploring large graphs with GPU-accelerated rendering and analytics via… | 97 | 2551 | active |
| scikit-learn-contrib/category_encoders A scikit-learn-contrib library providing a collection of sklearn-compatible transformers for encoding categorical variables into numeric fo… | 94 | 2499 | active |
| ArcticDB ArcticDB is a high-performance DataFrame database for time series and tick data, developed at Man Group for quantitative data science. It s… | 94 | 2493 | active |
| EntilZha/PyFunctional PyFunctional is a Python library for creating data pipelines using chained functional operators like map, filter, and reduce. Its API is in… | 28 | 2488 | stable |
| marin-community/marin Marin is an open-source Python framework and research program for training foundation models, covering the full pipeline from data curation… | 78 | 2428 | active |
| sodadata/soda-core Soda Core is a data quality and data contract verification engine that lets teams define data quality contracts in YAML and validate schema… | 99 | 2417 | active |
| julien-duponchelle/python-mysql-replication A pure Python library implementing the MySQL replication protocol on top of PyMySQL, letting applications stream binlog events such as inse… | 93 | 2414 | active |
| the-momentum/open-wearables Open Wearables is a self-hosted, MIT-licensed platform that unifies wearable health data from providers like Garmin, Whoop, Oura, Fitbit, S… | 83 | 2407 | active |
| apache/sedona Apache Sedona is a cluster computing framework for processing large-scale geospatial data across Spark, Flink, and Snowflake. It provides S… | 90 | 2390 | active |
| quarylabs/quary Quary is an open-source business intelligence tool for engineers that connects to databases, lets users write SQL to transform and document… | 81 | 2376 | active |
| moj-analytical-services/splink Splink is a Python library for fast, scalable probabilistic record linkage and entity resolution, based on the Fellegi-Sunter model with un… | 90 | 2364 | active |
| rsyslog/rsyslog Rsyslog is a high-performance open-source log ingestion and ETL engine written in C, originally a syslog daemon and now a modular pipeline … | 95 | 2329 | stable |
| warproxxx/poly_data A Python data pipeline that fetches, processes, and structures Polymarket trading data. It streams order events from the Polymarket CTF Exc… | 57 | 2304 | active |
| apache/gobblin Apache Gobblin is a distributed data integration framework for ingesting, replicating, organizing, and managing lifecycle of data across st… | 66 | 2270 | stable |
| ricklamers/gridstudio Grid studio is a web-based spreadsheet application with full integration of the Python runtime, backed by a Go spreadsheet engine. It provi… | 10 | 8817 | maintenance |
| timeplus-io/proton Timeplus Proton is a unified streaming SQL engine shipped as a single dependency-free C++ binary, built on the ClickHouse engine. It ingest… | 93 | 2247 | active |
| openmeterio/openmeter OpenMeter is an open-source metering and billing platform for AI, API, and DevOps products that ingests high-volume usage events, aggregate… | 95 | 2231 | active |
| rapidsai/cugraph cuGraph is NVIDIA's RAPIDS collection of GPU-accelerated graph analytics libraries, offering Python, C, and C++ APIs for building graphs an… | 94 | 2225 | active |
| ytsaurus/ytsaurus YTsaurus is an open-source, fault-tolerant big data platform combining distributed storage, MapReduce processing, a SQL query engine, and a… | 94 | 2200 | active |
| sequinstream/sequin Sequin is an open-source change data capture (CDC) platform for Postgres, deployed as a standalone Docker container that streams database c… | 63 | 2191 | active |
| tensorflow/tfx TensorFlow Extended (TFX) is an end-to-end, Google-production-scale platform for building and deploying production machine learning pipelin… | 91 | 2190 | stable |
| reugn/go-streams A lightweight stream processing library for Go offering a concise DSL to build declarative data pipelines from composable sources, flows, a… | 58 | 2170 | active |
| fugue-project/fugue Fugue is a Python library providing a unified interface for distributed computing, letting users run Python, Pandas, Polars, and SQL code o… | 82 | 2169 | active |
| log2timeline/plaso Plaso (log2timeline) is a Python-based engine for automatically creating super timelines from timestamped events found in logs and files on… | 87 | 2140 | active |
| apache/datafusion-ballista Apache DataFusion Ballista is a distributed query execution engine built on Apache DataFusion that parallelizes SQL and DataFrame workloads… | 98 | 2115 | active |
| warpstreamlabs/bento Bento is a high-performance, resilient stream processor written in Go that connects a wide range of sources and sinks (Kafka, Pub/Sub, Redi… | 91 | 2112 | active |
| peteromallet/dataclaw DataClaw is a CLI tool and macOS menu-bar app that parses coding-agent conversation history (Claude Code, Codex, Gemini CLI, etc.), redacts… | 72 | 2109 | active |
| cuemacro/findatapy findatapy is a Python library providing a unified high-level API for downloading market data from sources like Bloomberg, Refinitiv Eikon, … | 83 | 2108 | active |
| alibaba/otter Otter is Alibaba's distributed database synchronization system that parses database incremental logs (via Canal) to replicate data near-rea… | 23 | 8126 | maintenance |
| zorlan/skycaiji SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a … | 77 | 2089 | active |
| feldera/feldera Feldera is an incremental computation engine written in Rust that evaluates arbitrary SQL programs incrementally using DBSP theory, maintai… | 92 | 2064 | active |
| superglue-ai/superglue superglue is an open-source, self-hostable agentic integration platform that builds production-grade API integrations, data migrations, and… | 65 | 2053 | active |
| probberechts/soccerdata A Python library of scrapers that collect soccer data from popular websites like FBref, ESPN, WhoScored, Sofascore, SoFIFA, Understat, Club… | 93 | 2040 | active |
| elastic/elasticsearch-hadoop Elasticsearch-Hadoop (ES-Hadoop) is a Java connector library that integrates Elasticsearch real-time search and analytics with the Hadoop e… | 95 | 1972 | active |
| apache/cassandra-spark-connector The official Apache Spark connector for Apache Cassandra, letting Spark applications read Cassandra tables as RDDs and DataFrames, write th… | 31 | 1954 | active |
| tkfy920/qstock qstock is a Python library for personal quantitative investment research, providing modules for fetching financial market data (from Eastmo… | 37 | 1930 | active |
| NVIDIA/aistore AIStore (AIS) is a lightweight, distributed object storage stack built by NVIDIA specifically for AI workloads. It provides linear scalabil… | 95 | 1915 | active |
| johannfaouzi/pyts pyts is a Python package for time series classification that provides preprocessing, transformation, and utility tools along with implement… | 35 | 1876 | stable |
| neo4j-contrib/neo4j-apoc-procedures APOC (Awesome Procedures On Cypher) is a library of hundreds of procedures and functions that extend Neo4j's Cypher query language. It is p… | 93 | 1869 | active |
| pinterest/secor Secor is a Java service that persists Apache Kafka log messages to object stores such as Amazon S3, Google Cloud Storage, Azure Blob Storag… | 64 | 1856 | stable |
| davyxu/tabtoy Tabtoy is a high-performance command-line tool that exports spreadsheet data (Xlsx/CSV) into JSON, binary, and source code for Golang, C#, … | 23 | 1852 | active |
| buildship-ai/rowy Rowy is an open-source low-code backend platform for Firebase and Google Cloud that provides an Airtable-like spreadsheet UI for managing F… | 23 | 6834 | maintenance |
| dbt-labs/dbt-utils dbt-utils is an official dbt package providing reusable Jinja macros and generic data tests for dbt projects. It includes SQL generators (d… | 92 | 1792 | stable |
| apache/storm Apache Storm is a free and open-source distributed realtime computation system for processing unbounded streams of data, providing primitiv… | 97 | 6698 | maintenance |