Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: etl

671 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
nghuyong/WeiboSpider
A continuously maintained Python web scraping tool for Sina Weibo built on Scrapy and the new weibo.com API. It collects user profiles, pos…
734109active
DTStack/chunjun
ChunJun (formerly FlinkX) is a distributed, batch-and-stream data integration framework built on Apache Flink. It synchronizes and computes…
544099active
mahdibland/V2RayAggregator
A Python-based automation that aggregates free proxy nodes (Shadowsocks, SSR, Trojan, Vmess) from public sources, deduplicates them, and sp…
674008active
WikiExtractor/wikiextractor
WikiExtractor is a Python command-line tool that extracts and cleans plain text from Wikipedia database backup dumps (XML). It supports tem…
984001stable
ucbepic/docetl
DocETL is a Python library and CLI for building LLM-powered data processing and ETL pipelines over structured and unstructured data using d…
863995active
Rdatatable/data.table
data.table is an R package providing a high-performance, memory-efficient replacement for base R's data.frame with a concise indexing synta…
773913stable
Netflix/maestro
Netflix Maestro is a general-purpose workflow orchestrator providing workflow-as-a-service for data, ML, and software pipelines. It schedul…
693829active
dathere/qsv
qsv is a blazing-fast, Rust-based command-line data-wrangling toolkit with 50+ composable commands for querying, transforming, validating, …
983767active
jtablesaw/tablesaw
Tablesaw is a Java dataframe library for loading, cleaning, transforming, filtering, and summarizing tabular data, with descriptive statist…
863761active
awslabs/deequ
Deequ is a Scala library built on Apache Spark for defining 'unit tests for data' that measure data quality in large datasets. It computes …
933643active
ploomber/ploomber
Ploomber is a Python framework for building maintainable data pipelines from scripts and Jupyter notebooks, with iterative local developmen…
103622active
dtinit/data-transfer-project
An open-source framework providing common data models and adapter libraries that enable direct service-to-service transfer of user data bet…
673619active
waditu/tushare
TuShare is a Python library for crawling, cleaning, and storing historical and realtime financial data for China stocks and futures. It pro…
2315366maintenance
chrislusf/gleam
Gleam is a fast, efficient distributed map/reduce execution system written in pure Go, defining computations as DAG flows that can run stan…
753562active
noflo/noflo
NoFlo is a JavaScript implementation of Flow-Based Programming (FBP), where applications are defined as graphs of independent components co…
643552active
nextflow-io/nextflow
Nextflow is an open-source DSL and runtime for building scalable, portable, and reproducible data-driven computational pipelines, based on …
973475stable
ankane/pgsync
pgsync is a Ruby command-line tool that syncs data from one Postgres database to another, similar to pg_dump/pg_restore. It transfers table…
763468active
arpanghosh8453/garmin-grafana
A Dockerized Python application that fetches health and fitness data from Garmin Connect servers and stores it in a local InfluxDB time-ser…
793432active
xlwings/xlwings
xlwings is a BSD-licensed Python library that lets you call Python from Excel and automate Excel from Python, replacing VBA macros with Pyt…
993396active
apache/paimon
Apache Paimon is a lakehouse table format that combines lake format with LSM tree structure to support real-time streaming updates. It enab…
873384active
linuxfoundation/crowd.dev
LFX Community Data Platform (formerly crowd.dev) is a self-hostable platform that unifies community and developer engagement data from many…
673364active
hicccc77/WeFlow
WeFlow is a fully local desktop application for viewing, analyzing, and exporting WeChat chat history, including yearly data reports and vi…
5913959maintenance
django-import-export/django-import-export
A Django application and library for importing and exporting data between Django models and file formats like CSV, XLSX, JSON, YAML, and HT…
933334stable
lakehq/sail
Sail is an open-source, Rust-native multimodal compute engine that serves as a drop-in replacement for Apache Spark, unifying batch process…
933333active
huggingface/datatrove
DataTrove is a Python library from Hugging Face for processing, filtering, and deduplicating large-scale text data through customizable pip…
863308active
tcgoetz/GarminDB
A Python package and CLI that downloads and parses health and activity data from Garmin Connect, Garmin watches, FitBit CSV, and MS Health …
953277active
WeBankFinTech/DataSphereStudio
DataSphere Studio is a one-stop data application development and management portal from WeBank, built on the Linkis computing middleware. I…
503262active
SQLMesh/sqlmesh
SQLMesh is a scalable data transformation framework for running and deploying SQL or Python data models, backward compatible with dbt. It p…
933256active
PeerDB-io/peerdb
PeerDB is an open-source ETL/CDC tool that streams data from Postgres to data warehouses, queues, and storage engines, with an actively mai…
933252active
lakesoul-io/LakeSoul
LakeSoul is a cloud-native, real-time lakehouse framework with a Rust-native core providing ACID table format, concurrent upserts, incremen…
823247active
pydata/pandas-datareader
A Python library that extracts data from remote internet sources such as FRED, Fama/French, World Bank, OECD, and Eurostat directly into pa…
933238active
oxylabs/ai-crawler-py
A Python client library for Oxylabs AI Studio's AI-Crawler, a commercial service that crawls websites from a starting URL, uses natural lan…
613218active
duckdb/pg_duckdb
pg_duckdb is an official PostgreSQL extension that embeds DuckDB's columnar-vectorized analytics engine inside Postgres. It lets users run …
723210active
Wisser/Jailer
Jailer is a Java-based tool for database subsetting and relational data browsing. It extracts small, referentially intact slices of product…
993194active
webdataset/webdataset
A Python library providing a high-performance sequential I/O system based on tar-shard files for large-scale deep learning training, with s…
523169stable
omkarcloud/google-maps-scraper
A desktop application and API for scraping Google Maps business data, extracting 50+ data points including emails, phone numbers, social pr…
723130active
blockchain-etl/ethereum-etl
A Python CLI tool that extracts Ethereum blockchain data (blocks, transactions, token transfers, logs, traces, contracts) and exports it to…
523127active
apache/devlake
Apache DevLake is an open-source dev data platform that ingests, analyzes, and visualizes fragmented data from DevOps tools like GitHub, Gi…
933111active
alldatacenter/alldata
AllData is a definable data platform (数据中台) that integrates open-source components like DolphinScheduler, DataX, SeaTunnel, OpenMetadata, a…
833082active
TimelyDataflow/differential-dataflow
A Rust library implementing differential dataflow on top of timely dataflow, enabling data-parallel programs that efficiently process large…
932998active
spring-projects/spring-batch
Spring Batch is a lightweight, comprehensive batch framework for Java and Spring that enables development of robust batch applications. It …
952951stable
duckdb/ducklake
DuckLake is an integrated data lake and lakehouse catalog format that stores data as Parquet files and manages metadata in an ACID-complian…
652939active
cdhigh/KindleEar
KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep…
732865active
snakemake/snakemake
Snakemake is a Python-based workflow management system for creating reproducible and scalable data analyses. Workflows are defined via a re…
952853stable
numaproj/numaflow
Numaflow is a Kubernetes-native, serverless platform for running massively parallel data and streaming processing jobs. It lets developers …
982825active
datachain-ai/datachain
DataChain is a Python library that turns unstructured files in S3, GCS, and Azure into typed, versioned datasets queryable at warehouse spe…
862811active
dfinke/ImportExcel
A PowerShell module for importing and exporting Excel spreadsheets without requiring Excel to be installed. It supports creating tables, pi…
672733active
Ontos-AI/knowhere
Knowhere is a document ingestion and parsing platform that transforms unstructured documents (PDF, DOCX, XLSX, PPTX, images) into structure…
812673active
sfu-db/connector-x
ConnectorX is a Rust-based library with Python bindings that loads data from databases into DataFrames (pandas, Arrow, Polars) as fast and …
832646active
LLMQuant/quant-mind
QuantMind is an agent-native Python framework that refines raw financial information such as papers, news, and SEC filings into typed, time…
642630active
spotify/scio
Scio is a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. It provides a unified batch and strea…
962628active
OpenLineage/OpenLineage
OpenLineage is an open standard and framework for collecting data lineage metadata, defining a generic model of jobs, runs, and datasets wi…
972626active
dgunning/edgartools
EdgarTools is a Python library for accessing and parsing SEC EDGAR filings as structured, typed Python objects. It provides a clean API for…
942624active
meltano/meltano
Meltano is an open-source, declarative, code-first data integration engine and CLI for building and running ELT pipelines. It orchestrates …
972610stable
apache/hamilton
Apache Hamilton is a lightweight Python library for defining directed acyclic graphs (DAGs) of data transformations as plain, testable Pyth…
912575active
neilotoole/sq
sq is a command-line data wrangling tool that provides jq-style querying over SQL databases and document formats like CSV, Excel, and JSON.…
962555active
graphistry/pygraphistry
PyGraphistry is a Python library for loading, shaping, and visually exploring large graphs with GPU-accelerated rendering and analytics via…
972551active
scikit-learn-contrib/category_encoders
A scikit-learn-contrib library providing a collection of sklearn-compatible transformers for encoding categorical variables into numeric fo…
942499active
ArcticDB
ArcticDB is a high-performance DataFrame database for time series and tick data, developed at Man Group for quantitative data science. It s…
942493active
EntilZha/PyFunctional
PyFunctional is a Python library for creating data pipelines using chained functional operators like map, filter, and reduce. Its API is in…
282488stable
marin-community/marin
Marin is an open-source Python framework and research program for training foundation models, covering the full pipeline from data curation…
782428active
sodadata/soda-core
Soda Core is a data quality and data contract verification engine that lets teams define data quality contracts in YAML and validate schema…
992417active
julien-duponchelle/python-mysql-replication
A pure Python library implementing the MySQL replication protocol on top of PyMySQL, letting applications stream binlog events such as inse…
932414active
the-momentum/open-wearables
Open Wearables is a self-hosted, MIT-licensed platform that unifies wearable health data from providers like Garmin, Whoop, Oura, Fitbit, S…
832407active
apache/sedona
Apache Sedona is a cluster computing framework for processing large-scale geospatial data across Spark, Flink, and Snowflake. It provides S…
902390active
quarylabs/quary
Quary is an open-source business intelligence tool for engineers that connects to databases, lets users write SQL to transform and document…
812376active
moj-analytical-services/splink
Splink is a Python library for fast, scalable probabilistic record linkage and entity resolution, based on the Fellegi-Sunter model with un…
902364active
rsyslog/rsyslog
Rsyslog is a high-performance open-source log ingestion and ETL engine written in C, originally a syslog daemon and now a modular pipeline …
952329stable
warproxxx/poly_data
A Python data pipeline that fetches, processes, and structures Polymarket trading data. It streams order events from the Polymarket CTF Exc…
572304active
apache/gobblin
Apache Gobblin is a distributed data integration framework for ingesting, replicating, organizing, and managing lifecycle of data across st…
662270stable
ricklamers/gridstudio
Grid studio is a web-based spreadsheet application with full integration of the Python runtime, backed by a Go spreadsheet engine. It provi…
108817maintenance
timeplus-io/proton
Timeplus Proton is a unified streaming SQL engine shipped as a single dependency-free C++ binary, built on the ClickHouse engine. It ingest…
932247active
openmeterio/openmeter
OpenMeter is an open-source metering and billing platform for AI, API, and DevOps products that ingests high-volume usage events, aggregate…
952231active
rapidsai/cugraph
cuGraph is NVIDIA's RAPIDS collection of GPU-accelerated graph analytics libraries, offering Python, C, and C++ APIs for building graphs an…
942225active
ytsaurus/ytsaurus
YTsaurus is an open-source, fault-tolerant big data platform combining distributed storage, MapReduce processing, a SQL query engine, and a…
942200active
sequinstream/sequin
Sequin is an open-source change data capture (CDC) platform for Postgres, deployed as a standalone Docker container that streams database c…
632191active
tensorflow/tfx
TensorFlow Extended (TFX) is an end-to-end, Google-production-scale platform for building and deploying production machine learning pipelin…
912190stable
reugn/go-streams
A lightweight stream processing library for Go offering a concise DSL to build declarative data pipelines from composable sources, flows, a…
582170active
fugue-project/fugue
Fugue is a Python library providing a unified interface for distributed computing, letting users run Python, Pandas, Polars, and SQL code o…
822169active
log2timeline/plaso
Plaso (log2timeline) is a Python-based engine for automatically creating super timelines from timestamped events found in logs and files on…
872140active
apache/datafusion-ballista
Apache DataFusion Ballista is a distributed query execution engine built on Apache DataFusion that parallelizes SQL and DataFrame workloads…
982115active
warpstreamlabs/bento
Bento is a high-performance, resilient stream processor written in Go that connects a wide range of sources and sinks (Kafka, Pub/Sub, Redi…
912112active
peteromallet/dataclaw
DataClaw is a CLI tool and macOS menu-bar app that parses coding-agent conversation history (Claude Code, Codex, Gemini CLI, etc.), redacts…
722109active
cuemacro/findatapy
findatapy is a Python library providing a unified high-level API for downloading market data from sources like Bloomberg, Refinitiv Eikon, …
832108active
alibaba/otter
Otter is Alibaba's distributed database synchronization system that parses database incremental logs (via Canal) to replicate data near-rea…
238126maintenance
zorlan/skycaiji
SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a …
772089active
feldera/feldera
Feldera is an incremental computation engine written in Rust that evaluates arbitrary SQL programs incrementally using DBSP theory, maintai…
922064active
superglue-ai/superglue
superglue is an open-source, self-hostable agentic integration platform that builds production-grade API integrations, data migrations, and…
652053active
probberechts/soccerdata
A Python library of scrapers that collect soccer data from popular websites like FBref, ESPN, WhoScored, Sofascore, SoFIFA, Understat, Club…
932040active
elastic/elasticsearch-hadoop
Elasticsearch-Hadoop (ES-Hadoop) is a Java connector library that integrates Elasticsearch real-time search and analytics with the Hadoop e…
951972active
apache/cassandra-spark-connector
The official Apache Spark connector for Apache Cassandra, letting Spark applications read Cassandra tables as RDDs and DataFrames, write th…
311954active
tkfy920/qstock
qstock is a Python library for personal quantitative investment research, providing modules for fetching financial market data (from Eastmo…
371930active
NVIDIA/aistore
AIStore (AIS) is a lightweight, distributed object storage stack built by NVIDIA specifically for AI workloads. It provides linear scalabil…
951915active
johannfaouzi/pyts
pyts is a Python package for time series classification that provides preprocessing, transformation, and utility tools along with implement…
351876stable
neo4j-contrib/neo4j-apoc-procedures
APOC (Awesome Procedures On Cypher) is a library of hundreds of procedures and functions that extend Neo4j's Cypher query language. It is p…
931869active
pinterest/secor
Secor is a Java service that persists Apache Kafka log messages to object stores such as Amazon S3, Google Cloud Storage, Azure Blob Storag…
641856stable
davyxu/tabtoy
Tabtoy is a high-performance command-line tool that exports spreadsheet data (Xlsx/CSV) into JSON, binary, and source code for Golang, C#, …
231852active
buildship-ai/rowy
Rowy is an open-source low-code backend platform for Firebase and Google Cloud that provides an Airtable-like spreadsheet UI for managing F…
236834maintenance
dbt-labs/dbt-utils
dbt-utils is an official dbt package providing reusable Jinja macros and generic data tests for dbt projects. It includes SQL generators (d…
921792stable
apache/storm
Apache Storm is a free and open-source distributed realtime computation system for processing unbounded streams of data, providing primitiv…
976698maintenance

← prev page 2 / 7 next →