Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: etl

671 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
justlovemaki/CloudFlare-AI-Insight-Daily
A Cloudflare Workers-based content aggregation and generation platform that daily curates AI industry news, trending open-source projects, …
631779active
lf-edge/ekuiper
LF Edge eKuiper is a lightweight stream processing engine for IoT edge devices, offering SQL-based and graph-based rule engine for real-tim…
981731active
bespokelabsai/curator
Bespoke Curator is a Python library for bulk LLM inference and scalable synthetic data curation for post-training and structured data extra…
791719active
jldbc/pybaseball
pybaseball is a Python package that scrapes and retrieves current and historical baseball statistics from sources like MLB Statcast (Baseba…
501719active
4paradigm/OpenMLDB
OpenMLDB is an open-source machine learning database that acts as a feature platform, computing consistent features for both offline traini…
671710active
br-g/openf1
OpenF1 is a free, open-source API providing real-time and historical Formula 1 data, including lap timings, car telemetry, driver info, and…
711692active
osm2pgsql-dev/osm2pgsql
osm2pgsql is a command-line ETL tool that imports OpenStreetMap data in OSM, PBF, and O5M formats into a PostgreSQL/PostGIS database. It su…
861680stable
Bruin
Bruin is an end-to-end data platform whose CLI combines data ingestion (via its ingestr tool), SQL/Python/R transformations, and data quali…
911678active
Multiwoven/multiwoven
Multiwoven is an open-source Reverse ETL and data activation platform that syncs data from warehouses like Snowflake, BigQuery, Redshift, a…
901671active
mortada/fredapi
fredapi is a Python wrapper around the Federal Reserve Bank of St. Louis FRED and ALFRED web services for fetching economic data. It return…
521654active
skrub-data/skrub
skrub is a Python library that facilitates machine learning with dataframes, providing preprocessing, encoding, and wrangling tools for tab…
931650active
sunlabuiuc/PyHealth
PyHealth is an open-source Python toolkit for clinical deep learning, unifying healthcare datasets (EHRs, physiological signals, medical im…
911650active
huggingface/aisheets
Hugging Face AI Sheets is an open-source no-code web application for building, enriching, and transforming datasets using AI models. It can…
581642active
ariacom/Seal-Report
Seal Report & Task is an open-source (MIT) reporting and business intelligence framework written in C# for .NET, enabling users to build, p…
861632active
meta-llama/synthetic-data-kit
A Python CLI tool from Meta for generating high-quality synthetic datasets to fine-tune LLMs. It follows a four-stage pipeline (ingest, cre…
421632active
karlicoss/HPI
HPI (Human Programming Interface) is a Python package ('my') that unifies personal data from social networks, reading, health, location, me…
801628active
snorkel-team/snorkel
Snorkel is a Python library for programmatically building and managing training data using weak supervision, letting users write labeling f…
626001maintenance
OpenDota
OpenDota is an open-source Dota 2 data platform that provides a REST API and web UI for match and player statistics. It ingests data from V…
761627active
Snowflake-Labs/pg_lake
pg_lake is a set of PostgreSQL extensions that turn Postgres into a lakehouse engine for Apache Iceberg tables and raw data lake files in o…
821625active
lit26/finvizfinance
A Python library that scrapes and downloads financial data from the FinViz website, returning stock fundamentals, technicals, charts, news,…
791623active
nerevu/riko
riko is a pure Python stream processing library modeled after Yahoo! Pipes, combining reusable, configuration-driven modular pipes with syn…
881606active
cuducos/minha-receita
Minha Receita is an open-source project that consolidates Brazilian Federal Revenue (Receita Federal) CNPJ company data from scattered CSV …
101597active
paradigmxyz/cryo
cryo is a Rust-based CLI tool (also available as a Python package) that extracts Ethereum/EVM blockchain data into parquet, csv, , or Pytho…
191580active
getdozer/dozer
Dozer is a real-time data movement tool written in Rust that captures change data (CDC) from sources like Postgres, MySQL, Snowflake, and K…
231579active
quixio/quix-streams
Quix Streams is a pure Python framework for building real-time data pipelines and event-driven applications on Apache Kafka using a Streami…
971568active
dimitri/pgcopydb
pgcopydb is a command-line tool that automates copying a PostgreSQL database between two running servers, combining pg_dump/pg_restore for …
811549active
ilyakatz/data-migrate
A Ruby gem for Rails that lets you run data migrations alongside schema migrations. Data migrations live in db/data and are tracked in a da…
811549active
mosaicml/streaming
StreamingDataset is a Python library from MosaicML for fast, accurate streaming of training data from cloud object storage (S3, GCS, Azure,…
701547active
combust/mleap
MLeap is a serialization format (Bundle.ML) and portable execution engine for machine learning pipelines, implemented in Scala with Python …
951543active
finos/legend
Legend is an end-to-end data platform from FINOS (originated at Goldman Sachs) covering the full data lifecycle: modeling, curation, queryi…
771543active
tensorflow/gnn
TensorFlow GNN is a Python library for building Graph Neural Networks on TensorFlow, including a GraphTensor type for heterogeneous graphs,…
661541active
hi-primus/optimus
Optimus is a Python library for agile data preparation that provides a unified API over pandas, Dask, cuDF, Dask-cuDF, Vaex, and PySpark. I…
231536active
FinanceData/FinanceDataReader
FinanceDataReader is a Python library and CLI tool for reading financial data such as stock listings, stock prices, indexes, exchange rates…
691535active
uwdata/arquero
Arquero is a JavaScript library for query processing and transformation of array-backed, column-oriented data tables. It offers a fluent, d…
451534stable
BemiHQ/BemiDB
BemiDB is an open-source analytical data warehouse that combines built-in data source connectors (like Fivetran) with a Postgres-compatible…
551531active
pyper-dev/pyper
Pyper is a pure-Python library for building concurrent and parallel data pipelines using functional programming patterns. It unifies thread…
281518active
lqzhgood/Shmily
Shmily is a large umbrella project for exporting and archiving personal chat and communication records from QQ, WeChat, SMS, call logs, pho…
551507active
pyjanitor-devs/pyjanitor
pyjanitor is a Python library providing clean, readable APIs for data cleaning on pandas DataFrames, inspired by the R package janitor. It …
991498active
apache/inlong
Apache InLong is a one-stop, full-scenario integration framework for massive data, supporting data ingestion, synchronization, and subscrip…
901497stable
Vincentqyw/cv-arxiv-daily
An automated daily digest of computer vision and robotics arXiv papers (SLAM, SFM, visual localization, keypoint detection, image matching,…
771494active
Surfer-Org/Protocol
Surfer Protocol is an open-source framework for exporting personal data from platforms like Gmail, iMessages, Twitter, Notion, and ChatGPT.…
131475active
cube2222/octosql
OctoSQL is a Go CLI tool that lets you query, join, and transform data from multiple databases and file formats using SQL through a unified…
235264maintenance
sfirke/janitor
janitor is an R package with simple, user-friendly functions for examining and cleaning dirty data, such as formatting data.frame column na…
231455stable
dadoonet/fscrawler
FSCrawler is a Java-based file system crawler that indexes binary documents (PDF, MS Office, Open Office) into Elasticsearch, tracking new,…
881450active
apache/hop
Apache Hop is an open-source data and metadata orchestration platform for visually designing and running data integration pipelines and wor…
941447active
tidyverse/tidyr
tidyr is an R package from the tidyverse that provides tools for reshaping messy data into tidy form, including pivoting between long and w…
701438stable
datazip-inc/olake
OLake Go is a high-performance open-source EL (extract-load) engine written in Go that replicates databases (PostgreSQL, MySQL, MongoDB, Or…
841431active
winedarksea/AutoTS
AutoTS is a Python library for automated time series forecasting, offering dozens of sklearn-style models (statistical, ML, deep learning) …
951430active
wgzhao/Addax
Addax is a fast, extensible ETL tool for synchronizing data between heterogeneous SQL and NoSQL data sources, forked and evolved from Aliba…
981425active
toluaina/pgsync
PGSync is an open-source (MIT) change data capture tool that syncs PostgreSQL, MySQL, or MariaDB data to Elasticsearch or OpenSearch in rea…
981415active
fmind/mlops-python-package
A Python package template that provides a production-grade code base with MLOps best practices for building and deploying machine learning …
911415active
amphi-ai/amphi-etl
Amphi is a visual data preparation and ETL tool that lets users build data pipelines through a drag-and-drop interface while generating sta…
701401active
data-forge/data-forge-ts
Data-Forge is a TypeScript data transformation and analysis toolkit for JavaScript, inspired by Pandas and LINQ. It provides a DataFrame/Se…
691392active
SebastienZh/StockTradebyZ
A semi-automated stock screening application for China's A-share market that fetches daily K-line data via Tushare, applies quantitative pr…
511390active
thinh-vu/vnstock
Vnstock is an open-source Python library for extracting and analyzing Vietnam stock market data, returning data as pandas DataFrames via si…
901383active
okfn-brasil/querido-diario
Querido Diário is an open-source project by Open Knowledge Brasil that scrapes and aggregates Brazilian municipal official gazettes (diário…
751373active
locationtech/geotrellis
GeoTrellis is a Scala library and framework for high-performance reading, writing, and processing of geospatial raster and vector data. It …
841372stable
spatie/simple-excel
A PHP library for reading and writing simple Excel (xlsx) and CSV files with a fluent API. It uses generators and LazyCollections to keep m…
861368active
nextstrain/ncov
A Nextstrain build pipeline that analyzes SARS-CoV-2 viral genomes to understand their evolution and spread, producing phylogenetic visuali…
881363active
nf-core/rnaseq
nf-core/rnaseq is a Nextflow-based bioinformatics pipeline for analyzing RNA sequencing data, performing QC, trimming, (pseudo-)alignment w…
941357active
SpiderClub/weibospider
A distributed web crawler for Sina Weibo (Chinese microblogging platform) built with Python, Celery, and requests. It scrapes user profiles…
324793maintenance
duckdb/dbt-duckdb
dbt-duckdb is the dbt adapter plugin for DuckDB, an embedded OLAP database. It lets you run dbt data transformation pipelines in SQL or Pyt…
941341active
datavane/tis
TIS is an AI-native data integration platform built on DataX, Flink, and Flink-CDC that provides a visual Web-UI for zero-code batch and re…
901339active
subsquid/squid-sdk
Squid SDK is a TypeScript ETL toolkit for building blockchain data indexers, supporting Ethereum-like chains, Substrate, and Solana, with d…
981337active
rwynn/monstache
Monstache is a Go daemon that continuously syncs MongoDB collections into Elasticsearch (or OpenSearch) in realtime using change streams. I…
481331active
petl-developers/petl
petl is a general-purpose Python package for extracting, transforming and loading tables of data. It provides a lightweight, pure-Python to…
981316stable
GoogleCloudPlatform/DataflowTemplates
A collection of Google-provided Apache Beam pipeline templates for Google Cloud Dataflow that solve common in-cloud data tasks like import/…
951310active
Mojang/DataFixerUpper
DataFixerUpper is a Java library from Mojang for incrementally building, merging, and optimizing data transformations between schema versio…
591309active
arkflow-rs/arkflow
ArkFlow is a high-performance stream processing engine written in Rust on top of Tokio, connecting configurable inputs (Kafka, MQTT, HTTP, …
701302active
elixir-explorer/explorer
Explorer is an Elixir library providing series (one-dimensional) and dataframes (two-dimensional) for fast data exploration and manipulatio…
871291active
DTStack/Taier
Taier is a self-hosted distributed dispatching platform for big data that handles task submission, DAG-based scheduling, and operations/mai…
841284active
Nixtla/mlforecast
mlforecast is a Python framework for time series forecasting using any machine learning model with fit/predict methods, providing efficient…
961269active
meta-pytorch/data
TorchData is a PyTorch library providing scalable, performant data loading utilities, including StatefulDataLoader, a drop-in replacement f…
671263active
slothflowlabs/duckle
Duckle is an open-source ETL/ELT platform built on DuckDB that you self-host on your own servers or cloud, with a visual no-code/low-code c…
811262active
apache/datafusion-comet
Apache DataFusion Comet is a high-performance accelerator plugin for Apache Spark that keeps queries Arrow-native end-to-end, executing ope…
811262active
0xSero/ai-data-extraction
A Python toolkit of extraction scripts that pulls complete chat, agent, and code-context history from AI coding assistants like Cursor, Cla…
601257active
firecrawl/fire-enrich
Fire Enrich is an AI-powered data enrichment web application that transforms a list of email addresses into rich company datasets, includin…
391255active
astronomer/astronomer-cosmos
Astronomer Cosmos is an open-source Python library that converts dbt Core or dbt Fusion projects into Apache Airflow DAGs and task groups w…
981254active
sentinel-hub/eo-learn
eo-learn is a collection of open-source Python packages for accessing and processing spatio-temporal satellite imagery, built around modula…
511247active
zinggAI/zingg
Zingg is an ML-based tool for scalable master data management, entity resolution, identity resolution, and record deduplication. It runs on…
881243active
darold/ora2pg
Ora2Pg is a free Perl-based tool that migrates an Oracle database to PostgreSQL by scanning the source database and generating SQL scripts …
661231active
thetahealth/mirobody
Mirobody is an open-source, AI-native health data engine that collects readings from lab reports, wearables, and genomics, standardizes the…
841229active
kevwan/go-stash
go-stash is a high-performance, open-source server-side data processing pipeline written in Go that ingests data from Kafka, applies config…
601224active
marcboeker/gmail-to-sqlite
A Python CLI application that syncs Gmail messages into a local SQLite database, supporting incremental and full syncs with deletion detect…
561221active
skyplane-project/skyplane
Skyplane is a CLI tool for blazingly fast bulk data transfers between cloud object stores (AWS S3, Azure Blob, GCS, IBM COS) and local disk…
231216active
pytroll/satpy
Satpy is a Python library for reading, manipulating, and writing meteorological remote sensing data from earth-observing satellites. It sup…
861205active
apache/incubator-xtable
Apache XTable (incubating) is a cross-table converter that translates lakehouse table format metadata between Apache Hudi, Apache Iceberg, …
841205active
gityuanbao/share
A personal open-source repository whose main component is akshare_collector, a Python tool built on AKShare that collects Chinese financial…
641201active
xorbitsai/xorbits
Xorbits is an open-source distributed computing framework that scales Python data science and machine learning workloads from a laptop to l…
641199active
wireservice/agate
agate is a Python data analysis library optimized for humans instead of machines, offering a readable alternative to numpy and pandas. It p…
751198stable
marsupialtail/quokka
Quokka is a lightweight distributed dataflow/query engine written in Python, built on Ray, DuckDB, Polars, and Arrow, designed for stateful…
231192active
nextgenhealthcare/connect
Mirth Connect by NextGen Healthcare is an open-source healthcare integration engine that filters, transforms, extracts, and routes messages…
231190active
go-mysql-org/go-mysql-elasticsearch
A Go service that automatically syncs MySQL data into Elasticsearch, using mysqldump for initial load and binlog parsing for incremental up…
324148maintenance
xataio/pgstream
pgstream is an open-source change data capture (CDC) tool and Go library that replicates PostgreSQL data, including DDL/schema changes, to …
881186active
functime-org/functime
functime is a Python library for production-ready global forecasting and time-series feature extraction on large panel datasets, built on l…
701183active
machow/siuba
siuba is a Python library that ports R's dplyr syntax to pandas DataFrames and SQL databases, using pipe (>>) chaining and lazy expressions…
421183active
apache/amoro
Apache Amoro (incubating) is a Lakehouse management system built on open data lake formats like Iceberg, Paimon, and Mixed-Hive. It provide…
731171active
lhotse-speech/lhotse
Lhotse is a Python library for flexible, scalable preparation of multimodal (speech, audio, video, image, text) data for machine learning, …
891149active
Tavish9/any4lerobot
Any4LeRobot is a curated collection of Python utilities for the Hugging Face LeRobot robotics ecosystem, including dataset conversion scrip…
641140active
certtools/intelmq
IntelMQ is an open-source solution for IT security teams (CERTs, CSIRTs, SOCs) for collecting and processing security feeds using a message…
631133active

← prev page 3 / 7 next →