function: etl
671 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| databricks/lilac Lilac is an open-source tool for exploring, curating, and quality-controlling datasets used for training, fine-tuning, and monitoring LLMs.… | 10 | 1072 | maintenance |
| wj596/go-mysql-transfer A standalone Go application that syncs MySQL data in real time by listening to the MySQL binlog as a fake replica. It transforms row change… | 23 | 1066 | maintenance |
| SciRuby/daru Daru (Data Analysis in RUby) is a pure Ruby library providing DataFrame and Vector data structures for storing, analyzing, manipulating, an… | 23 | 1061 | maintenance |
| J535D165/recordlinkage A modular Python toolkit for record linkage and duplicate detection across datasets, built on pandas and numpy. It provides indexing (block… | 23 | 1060 | maintenance |
| wiseio/paratext ParaText is a C++ library for reading text files in parallel across multiple cores, with Python bindings and a fast CSV reader. It can load… | 32 | 1052 | maintenance |
| facebookresearch/cc_net CCNet is a Python pipeline from Facebook AI Research for downloading, deduplicating, and cleaning Common Crawl web data into high-quality m… | 10 | 1045 | maintenance |
| blaze/odo Odo is a Python library for migrating data between different containers, from in-memory structures like lists and pandas DataFrames to out-… | 23 | 1006 | maintenance |
| supabase/etl Supabase ETL is a high-performance Postgres logical replication engine written in Rust that performs initial table syncs and streams ongoin… | 70 | 2322 | experimental |
| rajasekarv/vega Vega (formerly native_spark) is a from-scratch reimplementation of Apache Spark's distributed data processing engine written in Rust. It ai… | 32 | 2229 | experimental |
| kkyon/botflow Botflow is a Python dataflow programming framework for building data pipelines using pipes and routes, with parallelism via coroutines and … | 62 | 1196 | experimental |
| FinHackCN/finhack FinHack is an extensible Python quantitative finance framework covering the full quant research workflow: data collection, factor computati… | 76 | 1145 | experimental |
| nucleuscloud/neosync Neosync is an open-source data security platform that detects PII, anonymizes production data, and generates synthetic data to sync across … | 10 | 4142 | abandoned |
| orchest/orchest Orchest is a self-hosted, browser-based tool for visually building and running data pipelines using notebooks and scripts in Python, R, or … | 10 | 4132 | abandoned |
| mholt/timeliner Timeliner is a Go CLI tool that archives your digital life—photos, social media posts, messages, and location history from services like Go… | 10 | 3558 | abandoned |
| mozilla-services/heka Heka is a Go-based tool for collecting data from many sources, performing in-flight processing, and delivering results to multiple destinat… | 10 | 3394 | abandoned |
| harismuneer/Ultimate-Social-Scrapers A collection of Python-based scraping tools that extract public data from Facebook, Instagram, and Twitter (X), including posts, media, fol… | 44 | 3151 | abandoned |
| airingursb/bilibili-user A Python web crawler that scrapes Bilibili user profiles (id, nickname, gender, avatar, level, birthday, location, etc.) and stores them in… | 32 | 3090 | abandoned |
| WeRead2Notion A Python tool that syncs WeChat Reading (WeRead) highlights and notes into a Notion database, typically run on a daily schedule via GitHub … | 61 | 2919 | abandoned |
| jprante/elasticsearch-jdbc A JDBC importer tool that fetches tabular data from relational databases via SQL queries and indexes it into Elasticsearch. It runs as a Ja… | 10 | 2809 | abandoned |
| FeatureBaseDB/featurebase FeatureBase (formerly Pilosa) is a distributed analytical database built entirely on bitmap indexes, offering low-latency SQL queries over … | 10 | 2521 | abandoned |
| instill-ai/instill-core Instill Core is a full-stack AI infrastructure tool for data, model, and pipeline orchestration, aimed at building AI-first applications wi… | 76 | 2321 | abandoned |
| chiphuyen/lazynlp A Python library for crawling, cleaning, and deduplicating web pages to build massive monolingual text datasets, suitable for training lang… | 23 | 2284 | abandoned |
| shyiko/mysql-binlog-connector-java A Java library for reading MySQL binary logs, both from binlog files and by tapping into the live MySQL replication stream. It supports GTI… | 23 | 2269 | abandoned |
| ByConity/ByConity ByConity is an open-source cloud-native distributed SQL data warehouse built on the ClickHouse 21.8 codebase, developed by ByteDance/Volcan… | 10 | 2239 | abandoned |
| dominictarr/event-stream EventStream is a Node.js library for composing pipelines of event streams using functional operations like map, filter, split, and merge. I… | 10 | 2175 | abandoned |
| minimaxir/facebook-page-post-scraper A Python script collection that scrapes all posts, reactions, and comments from public Facebook Pages and open Groups via the Facebook Grap… | 10 | 2135 | abandoned |
| twitter/summingbird Summingbird is a Scala library for writing MapReduce aggregation programs that look like native collection transformations and run on distr… | 10 | 2123 | abandoned |
| onyx-platform/onyx Onyx is a masterless, fault-tolerant distributed computation system written in pure Clojure that supports both batch and stream processing … | 10 | 2051 | abandoned |
| Tencent/TubeMQ TubeMQ was Tencent's high-performance message queue, donated to the Apache Software Foundation in 2019 and renamed Apache InLong. This repo… | 23 | 1996 | abandoned |
| torodb/stampede ToroDB Stampede replicates data from a MongoDB replica set into relational tables in PostgreSQL, translating the document structure into no… | 23 | 1757 | abandoned |
| ClimbsRocks/auto_ml auto_ml is a Python library for automated machine learning that handles feature engineering, model selection, and hyperparameter optimizati… | 23 | 1654 | abandoned |
| discoproject/disco Disco is an open-source distributed MapReduce framework written in Erlang with a Python job API, originally developed at Nokia Research Cen… | 32 | 1630 | abandoned |
| stripe-archive/mosql MoSQL is a Ruby tool that streams the contents of a MongoDB cluster into PostgreSQL, using an oplog tailer to keep the SQL mirror continuou… | 10 | 1616 | abandoned |
| mongodb/mongo-hadoop A Java library that lets MongoDB (or BSON backup files) serve as an input source or output destination for Hadoop MapReduce jobs, with inte… | 10 | 1552 | abandoned |
| fossasia/open-event-scraper A Python tool that parses Google spreadsheets and converts them into Open Event JSON format, originally built for FOSSASIA 2016 event data.… | 10 | 1517 | abandoned |
| qinxuye/cola Cola is a high-level distributed crawling framework in Python for scraping pages and extracting structured data from websites. The same cra… | 10 | 1499 | abandoned |
| compose/transporter Transporter is a Go CLI tool that syncs and transforms data between persistence engines such as MongoDB, PostgreSQL, Elasticsearch, MySQL, … | 10 | 1441 | abandoned |
| mesos/spark This is the original UC Berkeley AMPLab repository for Apache Spark, a fast cluster computing system supporting Java, Scala, and Python. Th… | 32 | 1418 | abandoned |
| DocNow/twarc twarc is a command line tool and Python library for collecting and archiving Twitter JSON data via the Twitter v1.1 and v2 APIs. It was dev… | 50 | 1394 | abandoned |
| Jefferson-Henrique/GetOldTweets-python A Python library that retrieves old tweets by mimicking the JSON calls Twitter Search makes in the browser, bypassing the official API's ti… | 32 | 1340 | abandoned |
| nytlabs/streamtools Streamtools is a graphical toolkit from The New York Times R&D Lab for exploring, analyzing, and modifying streams of real-time data. Users… | 23 | 1312 | abandoned |
| yahoo/CaffeOnSpark CaffeOnSpark is a Spark package that brings the Caffe deep learning framework to Hadoop and Spark clusters, enabling distributed neural net… | 10 | 1261 | abandoned |
| uber-archive/AthenaX AthenaX is a SQL-based streaming analytics platform open sourced by Uber, built on Apache Flink and Apache Calcite. It lets users run produ… | 10 | 1223 | abandoned |
| foolcage/fooltrader fooltrader is a Python quantitative analysis and trading framework that crawls, cleans, and structures market data (stocks, futures, forex,… | 23 | 1198 | abandoned |
| killrweather/killrweather KillrWeather is a Scala reference application demonstrating integration of Apache Spark Streaming, Apache Kafka, Apache Cassandra, and Akka… | 32 | 1178 | abandoned |
| richardwilly98/elasticsearch-river-mongodb An Elasticsearch plugin (river) that indexes MongoDB collections into Elasticsearch by tailing the MongoDB oplog of a replica set. It suppo… | 32 | 1121 | abandoned |
| nfldb A Python library (nflgame) providing an API to retrieve and read NFL Game Center JSON data, including real-time feeds useful for fantasy fo… | 23 | 1082 | abandoned |
| jorgecarleitao/arrow2 A Rust implementation of the Apache Arrow columnar in-memory format, supporting IO with Parquet, Avro, CSV, IPC, and Flight, plus compute k… | 10 | 1064 | abandoned |
| LockerProject/Locker Locker is an open-source personal data platform ('the me platform') that collects a user's data from various online services, apps, and dev… | 32 | 1063 | abandoned |
| littlstar/s3-lambda A Node.js library that lets you run lambda-style functions (forEach, map, reduce, filter) over S3 objects with concurrency control. It enab… | 32 | 1059 | abandoned |
| paulyoder/LinqToExcel A .NET library that lets developers query Excel spreadsheets and CSV files using LINQ syntax, mapping worksheet rows to strongly-typed obje… | 32 | 1058 | abandoned |
| klbostee/dumbo Dumbo is a Python module that makes writing and running Hadoop Streaming programs easy, providing a convenient Python API for MapReduce pro… | 32 | 1030 | abandoned |
| bcbio/bcbio-nextgen bcbio-nextgen is a validated, community-developed pipeline toolkit for high-throughput sequencing analysis, including variant calling, RNA-… | 23 | 1030 | abandoned |
| facebookarchive/bistro Bistro is a C++ distributed task scheduler framework from Facebook that schedules and runs distributed tasks, including data-parallel jobs,… | 10 | 1026 | abandoned |
| n8n n8n is a fair-code workflow automation platform with native AI capabilities, combining a visual canvas with custom JavaScript/Python code. … | 95 | 202527 | active |
| infiniflow/ragflow RAGFlow is an open-source Retrieval-Augmented Generation (RAG) engine that combines deep document understanding with agent orchestration to… | 93 | 89328 | active |
| unclecode/crawl4ai Crawl4AI is an open-source Python library that crawls websites with a headless browser and converts pages into clean, LLM-ready Markdown fo… | 89 | 79471 | active |
| scrapy/scrapy Scrapy is a fast, high-level web crawling and scraping framework for Python used to extract structured data from websites. It provides a fu… | 99 | 64048 | stable |
| sansan0/TrendRadar TrendRadar is a self-hostable AI-powered public opinion and trending-news monitor that aggregates hot topics across multiple platforms and … | 61 | 61855 | active |
| LlamaIndex LlamaIndex is a Python (with a TypeScript variant) framework for building LLM-powered applications, centered on ingesting, indexing, and qu… | 95 | 51883 | stable |
| microsoft/qlib Qlib is an AI-oriented quantitative investment platform from Microsoft that supports the full quant research workflow, from data processing… | 66 | 47960 | active |
| NaiboWang/EasySpider EasySpider is a free, open-source visual no-code web crawler and browser automation (RPA) tool where users design scraping tasks by clickin… | 87 | 44442 | active |
| ray-project/ray Ray is a unified open-source framework for scaling AI and Python applications, consisting of a core distributed runtime (tasks, actors, obj… | 99 | 43614 | stable |
| DuckDB DuckDB is an in-process analytical SQL database management system (OLAP) written in C++, designed to be fast, portable, and easy to embed. … | 97 | 40675 | stable |
| topoteretes/cognee Cognee is an open-source Python library and platform that gives AI agents persistent long-term memory by ingesting data in any format and b… | 91 | 30281 | active |
| PrefectHQ/prefect Prefect is an open-source workflow orchestration framework that turns Python functions into production-grade data pipelines using decorator… | 95 | 23693 | stable |
| recommenders-team/recommenders A Python library and collection of Jupyter notebooks with best practices for building, evaluating, and operationalizing recommendation syst… | 67 | 21864 | active |
| openobserve/openobserve OpenObserve is an open-source, cloud-native observability platform that unifies logs, metrics, traces, RUM, and LLM observability in a sing… | 93 | 21490 | active |
| xming521/WeClone WeClone is an end-to-end Python framework for creating a personal AI digital twin by fine-tuning large language models on your exported cha… | 80 | 18171 | active |
| windmill-labs/windmill Windmill is an open-source, self-hostable developer platform that turns scripts in Python, TypeScript, Go, Bash, SQL and other languages in… | 95 | 17687 | active |
| argoproj/argo-workflows Argo Workflows is an open-source, container-native workflow engine for orchestrating parallel jobs on Kubernetes, implemented as a Kubernet… | 99 | 16939 | stable |
| apache/doris Apache Doris is an open-source MPP-based real-time analytical database that delivers sub-second queries over massive datasets, combining a … | 99 | 15817 | stable |
| open-metadata/OpenMetadata OpenMetadata is an open-source metadata management platform that unifies data cataloging, lineage, data quality, governance, and business s… | 95 | 14986 | active |
| microsoft/RD-Agent RD-Agent is a Microsoft open-source framework that uses LLM-powered agents to automate research and development processes such as factor mi… | 75 | 14342 | active |
| apache/druid Apache Druid is a high-performance, distributed, real-time analytics database written in Java for fast OLAP-style slice-and-dice queries on… | 89 | 14045 | stable |
| trinodb/trino Trino is a fast, distributed ANSI SQL query engine for big data analytics, formerly known as PrestoSQL. It queries data in place across div… | 93 | 13183 | stable |
| StarRocks/starrocks StarRocks is a high-performance distributed OLAP database and query engine built on an MPP architecture with a fully vectorized execution e… | 95 | 12043 | active |
| dataelement/bisheng BISHENG is an open-source LLM DevOps (LLMOps) platform for building enterprise AI applications, offering GenAI workflow orchestration, RAG,… | 94 | 11911 | active |
| code4craft/webmagic WebMagic is a scalable web crawler framework for Java covering the full crawl lifecycle: downloading, URL management, content extraction (X… | 60 | 11678 | active |
| cocoindex-io/cocoindex CocoIndex is an open-source incremental data framework for AI, with a Rust engine and a declarative Python API that keeps sources (files, S… | 82 | 11409 | active |
| rerun-io/rerun Rerun is an open-source SDK and viewer for logging, storing, querying, and visualizing multi-rate multimodal data such as images, point clo… | 99 | 11362 | active |
| semantica-agi/semantica Semantica is a Python library providing graph-native infrastructure for building context graphs, knowledge graphs, and decision intelligenc… | 85 | 10920 | active |
| Netflix/metaflow Metaflow is a human-centric Python framework from Netflix for building, managing, and deploying real-life AI/ML and data science systems. I… | 95 | 10245 | stable |
| databendlabs/databend Databend is an open-source, cloud-native data warehouse built in Rust that runs entirely on object storage (S3, Azure, GCS). It unifies BI … | 93 | 9423 | active |
| spring-projects/spring-ai Spring AI is an application framework for AI engineering that provides Spring-friendly, portable abstractions for integrating AI models int… | 95 | 9358 | active |
| Deep Lake Deep Lake is an open-source database for AI that stores multimodal data (images, video, audio, text, embeddings, annotations) in a format o… | 77 | 9228 | active |
| apache/iceberg Apache Iceberg is a high-performance open table format for huge analytic datasets, bringing SQL table reliability to big data lakes. This r… | 90 | 9177 | stable |
| GoogleCloudPlatform/knowledge-catalog A collection of tools, agents, and samples for Google Cloud's Knowledge Catalog (formerly Dataplex), an AI-powered data catalog and metadat… | 58 | 8913 | active |
| fluent/fluent-bit Fluent Bit is a fast, lightweight telemetry agent written in C that collects, processes, and forwards logs, metrics, and traces from any so… | 99 | 8062 | stable |
| pythonstock/stock A full-stack stock analysis system built in Python using pandas, akshare, bokeh, stockstats, ta-lib, and tornado, deployed via Docker and d… | 59 | 7864 | active |
| adithya-s-k/omniparse OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into… | 49 | 7815 | active |
| andeya/pholcus Pholcus is a distributed, high-concurrency web crawler framework written in pure Go. It supports standalone, server, and client modes with … | 90 | 7577 | active |
| dbgate/dbgate DbGate is a cross-platform, open-source (no)SQL database manager for MySQL, PostgreSQL, SQL Server, MongoDB, SQLite, Redis and many other d… | 94 | 7285 | active |
| flyteorg/flyte Flyte is an open-source workflow orchestration platform for coordinating data, ML models, and AI agents at scale, authored in pure Python a… | 95 | 7245 | active |
| Alluxio/alluxio Alluxio is an open-source distributed caching and data orchestration platform that sits between compute frameworks (Spark, Presto, Trino, P… | 31 | 7231 | stable |
| Zipstack/unstract Unstract is an open-source, LLM-driven platform that extracts structured JSON data from unstructured documents such as PDFs, images, and sc… | 88 | 7172 | active |
| rocketride-org/rocketride-server RocketRide is an open-source AI pipeline engine with a high-throughput C++ runtime and 100+ Python-extensible nodes for building, debugging… | 80 | 7090 | active |
| deepseek-ai/DeepSpec DeepSpec is a full-stack Python codebase from DeepSeek for training and evaluating draft models used in speculative decoding of large langu… | 54 | 7041 | active |
| Hazelcast Hazelcast is a unified real-time data platform combining distributed stream processing with a fast, in-memory data store. It lets applicati… | 82 | 6604 | stable |
| apache/camel Apache Camel is an open-source integration framework offering 350+ connectors for databases, APIs, message brokers, and cloud services, wit… | 77 | 6300 | stable |