function: etl
671 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| baidu/bigflow Baidu Bigflow is a distributed computing framework offering simple, flexible Python APIs for writing data processing programs that can run … | 48 | 1131 | active |
| apache/iceberg-python PyIceberg is a Python implementation of the Apache Iceberg table format specification, providing programmatic access to Iceberg table metad… | 92 | 1124 | active |
| scratchdata/scratchdata Scratch Data is a self-hosted Go service that acts as a wrapper around analytics databases like DuckDB, ClickHouse, BigQuery, Snowflake, an… | 19 | 1120 | active |
| multiprocessio/dsq dsq is a Go command-line tool that lets you run SQL queries directly against data files such as JSON, CSV, TSV, Excel, and Parquet, powered… | 23 | 3866 | maintenance |
| bellingcat/auto-archiver A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supp… | 92 | 1109 | active |
| childe/gohangout Gohangout is a Logstash-like data processing tool written in Go that consumes events (commonly from Kafka), applies configurable filters, a… | 74 | 1103 | active |
| fraunhoferportugal/tsfel TSFEL is an open-source Python library for extracting features from time series signals across statistical, temporal, spectral, and fractal… | 53 | 1098 | active |
| apache/systemds Apache SystemDS is an open-source machine learning system covering the end-to-end data science lifecycle, from data cleaning and feature en… | 87 | 1097 | active |
| jf-tech/omniparser Omniparser is a native Go ETL library that streams input data in formats like CSV, JSON, XML, fixed-length text, and EDI/X12/EDIFACT, trans… | 26 | 1087 | active |
| intake/intake Intake is a lightweight Python package for describing data declaratively, gathering datasets into searchable catalogs, and loading data fro… | 72 | 1085 | active |
| stripe/sync-engine A self-hosted service that syncs your Stripe account data (customers, subscriptions, payments) into a Postgres database. Built in TypeScrip… | 91 | 1078 | active |
| OHDSI/CommonDataModel An R package and repository defining the OMOP Common Data Model (CDM), providing SQL DDL scripts and tools to generate or instantiate CDM t… | 88 | 1077 | active |
| linkedin/databus Databus is LinkedIn's source-agnostic distributed change data capture (CDC) system that reliably captures and streams primary data store ch… | 32 | 3679 | maintenance |
| hail-is/hail Hail is an open-source Python library for scalable exploration and analysis of genomic data, built on Spark, Scala, and C++ primitives for … | 91 | 1070 | active |
| lensesio/stream-reactor Stream Reactor is a collection of Apache 2.0 licensed Kafka Connect sinks and sources maintained by Lenses.io since 2016. It provides conne… | 95 | 1068 | active |
| Kotlin/dataframe Kotlin DataFrame is a typesafe in-memory structured data processing library for the JVM, reconciling Kotlin's static typing with dynamic da… | 90 | 1064 | active |
| unitedstates/congress A community-run Python toolkit that collects and converts official U.S. Congress data—bills, amendments, roll call votes, nominations, and … | 52 | 1060 | active |
| bigdatagenomics/adam ADAM is a genomics analysis platform built on Apache Spark that provides schemas and APIs for processing genomic data like reads, variants,… | 55 | 1057 | active |
| alibaba/Alink Alink is a machine learning algorithm platform built on Apache Flink, developed by Alibaba's PAI team. It provides a large library of batch… | 23 | 3611 | maintenance |
| datacontract/datacontract-cli An open-source Python CLI for working with data contracts using the Open Data Contract Standard (ODCS). It lints contracts, connects to dat… | 92 | 1053 | active |
| devlive-community/datacap DataCap is a self-hosted integrated platform for managing, transforming, integrating, and visualizing data across many data sources. It pro… | 73 | 1051 | active |
| turicas/brasil.io The backend of Brasil.IO, a platform that collects, cleans, and publishes Brazilian public open datasets in accessible formats. It automate… | 77 | 1044 | active |
| airbnb/chronon Chronon is an open-source data platform from Airbnb for computing, backfilling, and serving ML features. It handles batch and streaming fea… | 75 | 1043 | active |
| sqlparser/sqlflow_public SQLFlow is a tool that tracks column-level data lineage from SQL scripts across more than 20 major databases like Snowflake, Hive, Oracle, … | 69 | 1041 | active |
| twitter/scalding Scalding is a Scala library for specifying Hadoop MapReduce jobs, built on top of the Cascading framework. It provides a type-safe, functio… | 23 | 3523 | maintenance |
| aimeos/ai-woocommerce An Aimeos extension package that migrates WooCommerce (WordPress) shop data into an Aimeos e-commerce database. It transfers products, cate… | 63 | 1026 | active |
| liyupi/sql-generator A web application that generates structured, reusable SQL statements from JSON definitions, built with Vue3, TypeScript, Vite, Ant Design, … | 32 | 3426 | maintenance |
| CyrilFeng/karma Karma is a self-hosted data insight tool described as an 'executable mind map', built on Trino, that lets users configure SQL-based data so… | 32 | 1008 | active |
| fslaborg/Deedle Deedle is an easy-to-use .NET library for data frame and time series manipulation, designed for exploratory scientific programming in F# an… | 77 | 1006 | active |
| databricks/koalas Koalas implements the pandas DataFrame API on top of Apache Spark, letting data scientists use familiar pandas code on distributed big data… | 23 | 3372 | maintenance |
| go-gota/gota Gota is a Go library implementing DataFrames, Series, and data wrangling methods for tabular data manipulation. It supports loading data fr… | 10 | 3266 | maintenance |
| chrislusf/glow Glow is a pure-Go distributed computation library providing MapReduce-style flow APIs (Map, Filter, Reduce) that can run in parallel on a s… | 32 | 3218 | maintenance |
| ferventdesert/Hawk Hawk is a visual crawler and ETL IDE written in C#/WPF that lets users graphically scrape webpages, clean, transform, and store data withou… | 23 | 3213 | maintenance |
| blaze/blaze Blaze is a Python library that provides a NumPy/Pandas-like interface for querying data living in databases, files, and other computing sys… | 23 | 3189 | maintenance |
| libffcv/ffcv FFCV is a fast data loading system for PyTorch that dramatically increases data throughput in model training by replacing standard data loa… | 23 | 2993 | maintenance |
| scikit-learn-contrib/sklearn-pandas A Python library that bridges pandas DataFrames and scikit-learn by mapping DataFrame columns to sklearn transformations. Its DataFrameMapp… | 23 | 2842 | maintenance |
| mars-project/mars Mars is a tensor-based unified framework for large-scale data computation that provides NumPy-, pandas-, and scikit-learn-compatible APIs w… | 23 | 2742 | maintenance |
| douban/dpark DPark is a Python clone of Apache Spark, providing a MapReduce-like distributed computing framework that supports iterative computation. Jo… | 10 | 2663 | maintenance |
| Yelp/mrjob mrjob is a Python library for writing and running Hadoop Streaming MapReduce jobs, with support for Amazon EMR, Google Cloud Dataproc, self… | 66 | 2613 | maintenance |
| apache/logging-flume Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log-lik… | 77 | 2567 | maintenance |
| dbt-labs/dbt-core dbt Core is an open-source command-line tool that lets data analysts and engineers transform data in warehouses using SQL select statements… | 95 | 13696 | experimental |
| justinzm/gopup GoPUP is a Python library that provides convenient interfaces to a wide range of public Chinese data sources, including Baidu/Weibo/Google … | 23 | 2548 | maintenance |
| alibaba/yugong Yugong is a Java-based database migration and synchronization tool developed by Alibaba to move data from Oracle to MySQL/DRDS. It supports… | 23 | 2514 | maintenance |
| decaywood/XueQiuSuperSpider A Java 8 web scraping framework for collecting stock data from Xueqiu (Snowball) and other Chinese financial sites. It is built around comp… | 32 | 2433 | maintenance |
| salesforce/TransmogrifAI TransmogrifAI is an AutoML library written in Scala that runs on Apache Spark for building modular, reusable, strongly typed machine learni… | 61 | 2277 | maintenance |
| dotnet/spark .NET for Apache Spark provides high-performance C# and F# bindings for Apache Spark, exposing DataFrames, SparkSQL, and Structured Streamin… | 77 | 2098 | maintenance |
| mara/mara-pipelines Mara Pipelines is a lightweight, opinionated ETL/ELT framework for Python that sits between plain scripts and Apache Airflow. Pipelines are… | 23 | 2089 | maintenance |
| birdLark/LarkMidTable LarkMidTable (云雀) is a one-stop open-source data middleware platform covering metadata management, data warehouse development, data integra… | 32 | 2071 | maintenance |
| bytewax/bytewax Bytewax is a Python-first framework with a Rust-based distributed engine for stateful event and stream processing, inspired by Apache Flink… | 62 | 2047 | maintenance |
| Qihoo360/Quicksql Quicksql is a SQL analysis middleware that provides unified SQL querying across relational databases, non-relational databases, and SQL-les… | 23 | 2040 | maintenance |
| apache/drill Apache Drill is a distributed MPP (massively parallel processing) SQL query engine for self-describing data such as JSON, Parquet, and othe… | 67 | 2022 | maintenance |
| uber/petastorm Petastorm is a Python data access library from Uber that enables single-machine or distributed training and evaluation of deep learning mod… | 58 | 1891 | maintenance |
| h2oai/datatable datatable is a Python package for manipulating 2-dimensional tabular data structures (data frames), inspired by R's data.table. It emphasiz… | 65 | 1876 | maintenance |
| yougov/mongo-connector A Python-based pipeline tool that synchronizes data from a MongoDB cluster to target systems such as Solr, Elasticsearch, or another MongoD… | 23 | 1872 | maintenance |
| alvarobartt/investpy investpy is a Python package for retrieving recent and historical financial data from Investing.com, covering stocks, funds, ETFs, indices,… | 57 | 1851 | maintenance |
| byzer-org/byzer-lang Byzer (formerly MLSQL) is a low-code, SQL-like distributed programming language and engine for data pipelines, analytics, and AI, built aro… | 23 | 1835 | maintenance |
| embulk/embulk Embulk is an open-source, plugin-based parallel bulk data loader written in Java that transfers data between databases, storages, file form… | 62 | 1783 | maintenance |
| thbar/kiba Kiba is a Ruby ETL framework for defining and running data-processing pipelines with sources, transforms, and destinations via a Ruby DSL. … | 50 | 1775 | maintenance |
| vanus-labs/vanus Vanus is an open-source, cloud-native, serverless message queue with built-in event processing capabilities, written in Go. It connects Saa… | 23 | 1698 | maintenance |
| bytedance/bitsail BitSail is ByteDance's open-source distributed data integration engine built on Flink, supporting batch, streaming, and incremental data sy… | 10 | 1675 | maintenance |
| datamllab/tods TODS is a full-stack automated machine learning system for outlier detection on multivariate time-series data, developed by DATA Lab at Ric… | 32 | 1666 | maintenance |
| dashbitco/flow Flow is an Elixir library for expressing parallel computations on collections, built on top of GenStage. It provides an Enum/Stream-like AP… | 76 | 1621 | maintenance |
| python-bonobo/bonobo Bonobo is a lightweight extract-transform-load (ETL) framework for Python 3.5+ that streams data through a directed acyclic graph of plain … | 32 | 1613 | maintenance |
| cgarciae/pypeln Pypeln is a Python library for building concurrent, multi-stage data pipelines using processes, threads, or asyncio tasks through a single … | 23 | 1596 | maintenance |
| siddhi-io/siddhi Siddhi is a cloud-native stream processing and complex event processing (CEP) engine that executes Streaming SQL queries to capture events … | 72 | 1590 | maintenance |
| maxpumperla/elephas Elephas is a Python library that extends Keras to run distributed deep learning training on Apache Spark. It serializes Keras models on the… | 23 | 1577 | maintenance |
| san089/goodreads_etl_pipeline An end-to-end ETL pipeline that ingests Goodreads API data into an AWS S3 data lake, transforms it with Spark on EMR, and loads it into a R… | 32 | 1542 | maintenance |
| AxeldeRomblay/MLBox MLBox is a Python automated machine learning (AutoML) library that handles data preprocessing, feature selection, leak detection, and hyper… | 23 | 1536 | maintenance |
| Factual/drake Drake is a text-based data workflow tool that organizes command execution around data and its dependencies, functioning like GNU Make but d… | 23 | 1485 | maintenance |
| damklis/DataEngineeringProject An end-to-end data engineering project that scrapes news from RSS feeds via Airflow-scheduled Python scrapers and streams them through Kafk… | 32 | 1429 | maintenance |
| HASecuritySolutions/VulnWhisperer VulnWhisperer is a vulnerability management tool and report aggregator that pulls scan reports from scanners like Nessus, Qualys, OpenVAS, … | 10 | 1396 | maintenance |
| tensorflow/ecosystem A collection of templates and connectors for integrating TensorFlow with other open-source frameworks such as Kubernetes, Kubeflow, Spark, … | 10 | 1377 | maintenance |
| nathanmarz/cascalog Cascalog is a Clojure/Java data processing and querying library built on Hadoop, offering a high-level Datalog-like abstraction as a replac… | 32 | 1373 | maintenance |
| martinsbalodis/web-scraper-chrome-extension Web Scraper is a Chrome browser extension for extracting data from web pages without writing code. Users define sitemaps describing how to … | 32 | 1363 | maintenance |
| qiniu/logkit logkit is a Go-based server agent that collects logs and metrics from many sources (files, databases, Kafka, Redis, sockets, HTTP, SNMP) an… | 23 | 1359 | maintenance |
| lukasmartinelli/pgfutter pgfutter is a small Go CLI tool that imports CSV and line-delimited JSON files into PostgreSQL with a single command, automatically creatin… | 23 | 1345 | maintenance |
| ropensci/drake drake is an R package that acts as a Make-like pipeline toolkit for data science workflows, analyzing dependencies, skipping up-to-date ste… | 23 | 1342 | maintenance |
| treasure-data/digdag Digdag is an open-source workload automation system for building, running, scheduling, and monitoring task pipelines as directed acyclic gr… | 63 | 1333 | maintenance |
| lanyrd/mysql-postgresql-converter A Python script that converts MySQL database dumps into PostgreSQL-compatible SQL, originally built for Lanyrd's migration from MySQL to Po… | 32 | 1316 | maintenance |
| PatMartin/Dex Dex is a desktop data exploration and visualization application written in Java/Groovy on top of JavaFX. It combines ETL capabilities, mach… | 32 | 1315 | maintenance |
| python-streamz/streamz Streamz is a Python library for building pipelines that manage continuous streams of real-time data. It supports complex pipelines with bra… | 66 | 1303 | maintenance |
| rocketlaunchr/dataframe-go dataframe-go is a lightweight Go library providing DataFrame and Series data structures for statistics, machine-learning, and data manipula… | 32 | 1292 | maintenance |
| microsoft/Trill Trill is a high-performance, single-node, one-pass in-memory streaming analytics engine from Microsoft Research, built on a temporal data a… | 32 | 1271 | maintenance |
| google-research/deduplicate-text-datasets A Rust implementation of ExactSubstr deduplication for language model training datasets, with Python scripts for running deduplication and … | 10 | 1270 | maintenance |
| ycjuan/kaggle-2014-criteo The winning '3 Idiots' solution code for the 2014 Kaggle Criteo Display Advertising Challenge, built around field-aware factorization machi… | 10 | 1254 | maintenance |
| 2ndQuadrant/pglogical pglogical is a PostgreSQL extension providing logical streaming replication using a publish/subscribe model, written in C. It supports sele… | 84 | 1238 | maintenance |
| AlgoTraders/stock-analysis-engine A distributed stock analysis and backtesting framework that ingests automated pricing data from IEX Cloud, Tradier, and FinViz and runs tho… | 32 | 1238 | maintenance |
| calogica/dbt-expectations A dbt extension package that ports Great Expectations-style data quality tests as dbt test macros. It lets data teams run expectation tests… | 23 | 1234 | maintenance |
| ultralytics/JSON2YOLO A legacy Python toolkit that converts JSON-format annotation datasets (COCO, LabelMe, Labelbox, VoTT, INFOLKS, ATH) into the YOLO format fo… | 67 | 1231 | maintenance |
| BriData/DBus DBus is a Java-based data bus platform that captures incremental changes from databases (via log-based CDC) and log sources in a non-intrus… | 23 | 1214 | maintenance |
| twosigma/flint Flint is Two Sigma's open-source time series library for Apache Spark, built around a time series aware TimeSeriesRDD data structure. It pr… | 32 | 1177 | maintenance |
| apache/griffin Apache Griffin is a model-driven data quality service platform for defining, executing, and reporting data quality measures across multiple… | 10 | 1172 | maintenance |
| NVIDIA-Merlin/NVTabular NVTabular is a GPU-accelerated feature engineering and preprocessing library for tabular data, built to handle terabyte-scale datasets for … | 60 | 1150 | maintenance |
| lensacom/sparkit-learn Sparkit-learn provides scikit-learn's API and functionality on top of PySpark, operating on distributed RDDs of numpy arrays and sparse mat… | 23 | 1150 | maintenance |
| twitter/elephant-bird Twitter's Java library of LZO, Thrift, and Protocol Buffer-related Hadoop InputFormats, Pig LoadFuncs, Hive SerDes, and HBase utilities. It… | 23 | 1133 | maintenance |
| oeljeklaus-you/UserActionAnalyzePlatform A big data platform for e-commerce user behavior analysis built on Spark (Core, SQL, Streaming) with Java. It provides four analysis module… | 32 | 1127 | maintenance |
| ucarGroup/DataLink DataLink is a distributed, extensible data exchange platform for real-time incremental and offline full synchronization between heterogeneo… | 23 | 1120 | maintenance |
| Teradata/kylo Kylo is an open-source enterprise data lake management platform for self-service data ingest and preparation, with integrated metadata mana… | 32 | 1112 | maintenance |
| rhiever/datacleaner A Python package and command-line tool that automatically cleans pandas DataFrames, handling missing values and encoding categorical variab… | 23 | 1079 | maintenance |
| markwk/qs_ledger A personal data aggregator and analysis toolkit for quantified self enthusiasts, written in Python and distributed as Jupyter Notebooks. It… | 32 | 1073 | maintenance |