Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: etl

671 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
baidu/bigflow
Baidu Bigflow is a distributed computing framework offering simple, flexible Python APIs for writing data processing programs that can run …
481131active
apache/iceberg-python
PyIceberg is a Python implementation of the Apache Iceberg table format specification, providing programmatic access to Iceberg table metad…
921124active
scratchdata/scratchdata
Scratch Data is a self-hosted Go service that acts as a wrapper around analytics databases like DuckDB, ClickHouse, BigQuery, Snowflake, an…
191120active
multiprocessio/dsq
dsq is a Go command-line tool that lets you run SQL queries directly against data files such as JSON, CSV, TSV, Excel, and Parquet, powered…
233866maintenance
bellingcat/auto-archiver
A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supp…
921109active
childe/gohangout
Gohangout is a Logstash-like data processing tool written in Go that consumes events (commonly from Kafka), applies configurable filters, a…
741103active
fraunhoferportugal/tsfel
TSFEL is an open-source Python library for extracting features from time series signals across statistical, temporal, spectral, and fractal…
531098active
apache/systemds
Apache SystemDS is an open-source machine learning system covering the end-to-end data science lifecycle, from data cleaning and feature en…
871097active
jf-tech/omniparser
Omniparser is a native Go ETL library that streams input data in formats like CSV, JSON, XML, fixed-length text, and EDI/X12/EDIFACT, trans…
261087active
intake/intake
Intake is a lightweight Python package for describing data declaratively, gathering datasets into searchable catalogs, and loading data fro…
721085active
stripe/sync-engine
A self-hosted service that syncs your Stripe account data (customers, subscriptions, payments) into a Postgres database. Built in TypeScrip…
911078active
OHDSI/CommonDataModel
An R package and repository defining the OMOP Common Data Model (CDM), providing SQL DDL scripts and tools to generate or instantiate CDM t…
881077active
linkedin/databus
Databus is LinkedIn's source-agnostic distributed change data capture (CDC) system that reliably captures and streams primary data store ch…
323679maintenance
hail-is/hail
Hail is an open-source Python library for scalable exploration and analysis of genomic data, built on Spark, Scala, and C++ primitives for …
911070active
lensesio/stream-reactor
Stream Reactor is a collection of Apache 2.0 licensed Kafka Connect sinks and sources maintained by Lenses.io since 2016. It provides conne…
951068active
Kotlin/dataframe
Kotlin DataFrame is a typesafe in-memory structured data processing library for the JVM, reconciling Kotlin's static typing with dynamic da…
901064active
unitedstates/congress
A community-run Python toolkit that collects and converts official U.S. Congress data—bills, amendments, roll call votes, nominations, and …
521060active
bigdatagenomics/adam
ADAM is a genomics analysis platform built on Apache Spark that provides schemas and APIs for processing genomic data like reads, variants,…
551057active
alibaba/Alink
Alink is a machine learning algorithm platform built on Apache Flink, developed by Alibaba's PAI team. It provides a large library of batch…
233611maintenance
datacontract/datacontract-cli
An open-source Python CLI for working with data contracts using the Open Data Contract Standard (ODCS). It lints contracts, connects to dat…
921053active
devlive-community/datacap
DataCap is a self-hosted integrated platform for managing, transforming, integrating, and visualizing data across many data sources. It pro…
731051active
turicas/brasil.io
The backend of Brasil.IO, a platform that collects, cleans, and publishes Brazilian public open datasets in accessible formats. It automate…
771044active
airbnb/chronon
Chronon is an open-source data platform from Airbnb for computing, backfilling, and serving ML features. It handles batch and streaming fea…
751043active
sqlparser/sqlflow_public
SQLFlow is a tool that tracks column-level data lineage from SQL scripts across more than 20 major databases like Snowflake, Hive, Oracle, …
691041active
twitter/scalding
Scalding is a Scala library for specifying Hadoop MapReduce jobs, built on top of the Cascading framework. It provides a type-safe, functio…
233523maintenance
aimeos/ai-woocommerce
An Aimeos extension package that migrates WooCommerce (WordPress) shop data into an Aimeos e-commerce database. It transfers products, cate…
631026active
liyupi/sql-generator
A web application that generates structured, reusable SQL statements from JSON definitions, built with Vue3, TypeScript, Vite, Ant Design, …
323426maintenance
CyrilFeng/karma
Karma is a self-hosted data insight tool described as an 'executable mind map', built on Trino, that lets users configure SQL-based data so…
321008active
fslaborg/Deedle
Deedle is an easy-to-use .NET library for data frame and time series manipulation, designed for exploratory scientific programming in F# an…
771006active
databricks/koalas
Koalas implements the pandas DataFrame API on top of Apache Spark, letting data scientists use familiar pandas code on distributed big data…
233372maintenance
go-gota/gota
Gota is a Go library implementing DataFrames, Series, and data wrangling methods for tabular data manipulation. It supports loading data fr…
103266maintenance
chrislusf/glow
Glow is a pure-Go distributed computation library providing MapReduce-style flow APIs (Map, Filter, Reduce) that can run in parallel on a s…
323218maintenance
ferventdesert/Hawk
Hawk is a visual crawler and ETL IDE written in C#/WPF that lets users graphically scrape webpages, clean, transform, and store data withou…
233213maintenance
blaze/blaze
Blaze is a Python library that provides a NumPy/Pandas-like interface for querying data living in databases, files, and other computing sys…
233189maintenance
libffcv/ffcv
FFCV is a fast data loading system for PyTorch that dramatically increases data throughput in model training by replacing standard data loa…
232993maintenance
scikit-learn-contrib/sklearn-pandas
A Python library that bridges pandas DataFrames and scikit-learn by mapping DataFrame columns to sklearn transformations. Its DataFrameMapp…
232842maintenance
mars-project/mars
Mars is a tensor-based unified framework for large-scale data computation that provides NumPy-, pandas-, and scikit-learn-compatible APIs w…
232742maintenance
douban/dpark
DPark is a Python clone of Apache Spark, providing a MapReduce-like distributed computing framework that supports iterative computation. Jo…
102663maintenance
Yelp/mrjob
mrjob is a Python library for writing and running Hadoop Streaming MapReduce jobs, with support for Amazon EMR, Google Cloud Dataproc, self…
662613maintenance
apache/logging-flume
Apache Flume is a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large amounts of log-lik…
772567maintenance
dbt-labs/dbt-core
dbt Core is an open-source command-line tool that lets data analysts and engineers transform data in warehouses using SQL select statements…
9513696experimental
justinzm/gopup
GoPUP is a Python library that provides convenient interfaces to a wide range of public Chinese data sources, including Baidu/Weibo/Google …
232548maintenance
alibaba/yugong
Yugong is a Java-based database migration and synchronization tool developed by Alibaba to move data from Oracle to MySQL/DRDS. It supports…
232514maintenance
decaywood/XueQiuSuperSpider
A Java 8 web scraping framework for collecting stock data from Xueqiu (Snowball) and other Chinese financial sites. It is built around comp…
322433maintenance
salesforce/TransmogrifAI
TransmogrifAI is an AutoML library written in Scala that runs on Apache Spark for building modular, reusable, strongly typed machine learni…
612277maintenance
dotnet/spark
.NET for Apache Spark provides high-performance C# and F# bindings for Apache Spark, exposing DataFrames, SparkSQL, and Structured Streamin…
772098maintenance
mara/mara-pipelines
Mara Pipelines is a lightweight, opinionated ETL/ELT framework for Python that sits between plain scripts and Apache Airflow. Pipelines are…
232089maintenance
birdLark/LarkMidTable
LarkMidTable (云雀) is a one-stop open-source data middleware platform covering metadata management, data warehouse development, data integra…
322071maintenance
bytewax/bytewax
Bytewax is a Python-first framework with a Rust-based distributed engine for stateful event and stream processing, inspired by Apache Flink…
622047maintenance
Qihoo360/Quicksql
Quicksql is a SQL analysis middleware that provides unified SQL querying across relational databases, non-relational databases, and SQL-les…
232040maintenance
apache/drill
Apache Drill is a distributed MPP (massively parallel processing) SQL query engine for self-describing data such as JSON, Parquet, and othe…
672022maintenance
uber/petastorm
Petastorm is a Python data access library from Uber that enables single-machine or distributed training and evaluation of deep learning mod…
581891maintenance
h2oai/datatable
datatable is a Python package for manipulating 2-dimensional tabular data structures (data frames), inspired by R's data.table. It emphasiz…
651876maintenance
yougov/mongo-connector
A Python-based pipeline tool that synchronizes data from a MongoDB cluster to target systems such as Solr, Elasticsearch, or another MongoD…
231872maintenance
alvarobartt/investpy
investpy is a Python package for retrieving recent and historical financial data from Investing.com, covering stocks, funds, ETFs, indices,…
571851maintenance
byzer-org/byzer-lang
Byzer (formerly MLSQL) is a low-code, SQL-like distributed programming language and engine for data pipelines, analytics, and AI, built aro…
231835maintenance
embulk/embulk
Embulk is an open-source, plugin-based parallel bulk data loader written in Java that transfers data between databases, storages, file form…
621783maintenance
thbar/kiba
Kiba is a Ruby ETL framework for defining and running data-processing pipelines with sources, transforms, and destinations via a Ruby DSL. …
501775maintenance
vanus-labs/vanus
Vanus is an open-source, cloud-native, serverless message queue with built-in event processing capabilities, written in Go. It connects Saa…
231698maintenance
bytedance/bitsail
BitSail is ByteDance's open-source distributed data integration engine built on Flink, supporting batch, streaming, and incremental data sy…
101675maintenance
datamllab/tods
TODS is a full-stack automated machine learning system for outlier detection on multivariate time-series data, developed by DATA Lab at Ric…
321666maintenance
dashbitco/flow
Flow is an Elixir library for expressing parallel computations on collections, built on top of GenStage. It provides an Enum/Stream-like AP…
761621maintenance
python-bonobo/bonobo
Bonobo is a lightweight extract-transform-load (ETL) framework for Python 3.5+ that streams data through a directed acyclic graph of plain …
321613maintenance
cgarciae/pypeln
Pypeln is a Python library for building concurrent, multi-stage data pipelines using processes, threads, or asyncio tasks through a single …
231596maintenance
siddhi-io/siddhi
Siddhi is a cloud-native stream processing and complex event processing (CEP) engine that executes Streaming SQL queries to capture events …
721590maintenance
maxpumperla/elephas
Elephas is a Python library that extends Keras to run distributed deep learning training on Apache Spark. It serializes Keras models on the…
231577maintenance
san089/goodreads_etl_pipeline
An end-to-end ETL pipeline that ingests Goodreads API data into an AWS S3 data lake, transforms it with Spark on EMR, and loads it into a R…
321542maintenance
AxeldeRomblay/MLBox
MLBox is a Python automated machine learning (AutoML) library that handles data preprocessing, feature selection, leak detection, and hyper…
231536maintenance
Factual/drake
Drake is a text-based data workflow tool that organizes command execution around data and its dependencies, functioning like GNU Make but d…
231485maintenance
damklis/DataEngineeringProject
An end-to-end data engineering project that scrapes news from RSS feeds via Airflow-scheduled Python scrapers and streams them through Kafk…
321429maintenance
HASecuritySolutions/VulnWhisperer
VulnWhisperer is a vulnerability management tool and report aggregator that pulls scan reports from scanners like Nessus, Qualys, OpenVAS, …
101396maintenance
tensorflow/ecosystem
A collection of templates and connectors for integrating TensorFlow with other open-source frameworks such as Kubernetes, Kubeflow, Spark, …
101377maintenance
nathanmarz/cascalog
Cascalog is a Clojure/Java data processing and querying library built on Hadoop, offering a high-level Datalog-like abstraction as a replac…
321373maintenance
martinsbalodis/web-scraper-chrome-extension
Web Scraper is a Chrome browser extension for extracting data from web pages without writing code. Users define sitemaps describing how to …
321363maintenance
qiniu/logkit
logkit is a Go-based server agent that collects logs and metrics from many sources (files, databases, Kafka, Redis, sockets, HTTP, SNMP) an…
231359maintenance
lukasmartinelli/pgfutter
pgfutter is a small Go CLI tool that imports CSV and line-delimited JSON files into PostgreSQL with a single command, automatically creatin…
231345maintenance
ropensci/drake
drake is an R package that acts as a Make-like pipeline toolkit for data science workflows, analyzing dependencies, skipping up-to-date ste…
231342maintenance
treasure-data/digdag
Digdag is an open-source workload automation system for building, running, scheduling, and monitoring task pipelines as directed acyclic gr…
631333maintenance
lanyrd/mysql-postgresql-converter
A Python script that converts MySQL database dumps into PostgreSQL-compatible SQL, originally built for Lanyrd's migration from MySQL to Po…
321316maintenance
PatMartin/Dex
Dex is a desktop data exploration and visualization application written in Java/Groovy on top of JavaFX. It combines ETL capabilities, mach…
321315maintenance
python-streamz/streamz
Streamz is a Python library for building pipelines that manage continuous streams of real-time data. It supports complex pipelines with bra…
661303maintenance
rocketlaunchr/dataframe-go
dataframe-go is a lightweight Go library providing DataFrame and Series data structures for statistics, machine-learning, and data manipula…
321292maintenance
microsoft/Trill
Trill is a high-performance, single-node, one-pass in-memory streaming analytics engine from Microsoft Research, built on a temporal data a…
321271maintenance
google-research/deduplicate-text-datasets
A Rust implementation of ExactSubstr deduplication for language model training datasets, with Python scripts for running deduplication and …
101270maintenance
ycjuan/kaggle-2014-criteo
The winning '3 Idiots' solution code for the 2014 Kaggle Criteo Display Advertising Challenge, built around field-aware factorization machi…
101254maintenance
2ndQuadrant/pglogical
pglogical is a PostgreSQL extension providing logical streaming replication using a publish/subscribe model, written in C. It supports sele…
841238maintenance
AlgoTraders/stock-analysis-engine
A distributed stock analysis and backtesting framework that ingests automated pricing data from IEX Cloud, Tradier, and FinViz and runs tho…
321238maintenance
calogica/dbt-expectations
A dbt extension package that ports Great Expectations-style data quality tests as dbt test macros. It lets data teams run expectation tests…
231234maintenance
ultralytics/JSON2YOLO
A legacy Python toolkit that converts JSON-format annotation datasets (COCO, LabelMe, Labelbox, VoTT, INFOLKS, ATH) into the YOLO format fo…
671231maintenance
BriData/DBus
DBus is a Java-based data bus platform that captures incremental changes from databases (via log-based CDC) and log sources in a non-intrus…
231214maintenance
twosigma/flint
Flint is Two Sigma's open-source time series library for Apache Spark, built around a time series aware TimeSeriesRDD data structure. It pr…
321177maintenance
apache/griffin
Apache Griffin is a model-driven data quality service platform for defining, executing, and reporting data quality measures across multiple…
101172maintenance
NVIDIA-Merlin/NVTabular
NVTabular is a GPU-accelerated feature engineering and preprocessing library for tabular data, built to handle terabyte-scale datasets for …
601150maintenance
lensacom/sparkit-learn
Sparkit-learn provides scikit-learn's API and functionality on top of PySpark, operating on distributed RDDs of numpy arrays and sparse mat…
231150maintenance
twitter/elephant-bird
Twitter's Java library of LZO, Thrift, and Protocol Buffer-related Hadoop InputFormats, Pig LoadFuncs, Hive SerDes, and HBase utilities. It…
231133maintenance
oeljeklaus-you/UserActionAnalyzePlatform
A big data platform for e-commerce user behavior analysis built on Spark (Core, SQL, Streaming) with Java. It provides four analysis module…
321127maintenance
ucarGroup/DataLink
DataLink is a distributed, extensible data exchange platform for real-time incremental and offline full synchronization between heterogeneo…
231120maintenance
Teradata/kylo
Kylo is an open-source enterprise data lake management platform for self-service data ingest and preparation, with integrated metadata mana…
321112maintenance
rhiever/datacleaner
A Python package and command-line tool that automatically cleans pandas DataFrames, handling missing values and encoding categorical variab…
231079maintenance
markwk/qs_ledger
A personal data aggregator and analysis toolkit for quantified self enthusiasts, written in Python and distributed as Jupyter Notebooks. It…
321073maintenance

← prev page 4 / 7 next →