Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: etl

671 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
databricks/lilac
Lilac is an open-source tool for exploring, curating, and quality-controlling datasets used for training, fine-tuning, and monitoring LLMs.…
101072maintenance
wj596/go-mysql-transfer
A standalone Go application that syncs MySQL data in real time by listening to the MySQL binlog as a fake replica. It transforms row change…
231066maintenance
SciRuby/daru
Daru (Data Analysis in RUby) is a pure Ruby library providing DataFrame and Vector data structures for storing, analyzing, manipulating, an…
231061maintenance
J535D165/recordlinkage
A modular Python toolkit for record linkage and duplicate detection across datasets, built on pandas and numpy. It provides indexing (block…
231060maintenance
wiseio/paratext
ParaText is a C++ library for reading text files in parallel across multiple cores, with Python bindings and a fast CSV reader. It can load…
321052maintenance
facebookresearch/cc_net
CCNet is a Python pipeline from Facebook AI Research for downloading, deduplicating, and cleaning Common Crawl web data into high-quality m…
101045maintenance
blaze/odo
Odo is a Python library for migrating data between different containers, from in-memory structures like lists and pandas DataFrames to out-…
231006maintenance
supabase/etl
Supabase ETL is a high-performance Postgres logical replication engine written in Rust that performs initial table syncs and streams ongoin…
702322experimental
rajasekarv/vega
Vega (formerly native_spark) is a from-scratch reimplementation of Apache Spark's distributed data processing engine written in Rust. It ai…
322229experimental
kkyon/botflow
Botflow is a Python dataflow programming framework for building data pipelines using pipes and routes, with parallelism via coroutines and …
621196experimental
FinHackCN/finhack
FinHack is an extensible Python quantitative finance framework covering the full quant research workflow: data collection, factor computati…
761145experimental
nucleuscloud/neosync
Neosync is an open-source data security platform that detects PII, anonymizes production data, and generates synthetic data to sync across …
104142abandoned
orchest/orchest
Orchest is a self-hosted, browser-based tool for visually building and running data pipelines using notebooks and scripts in Python, R, or …
104132abandoned
mholt/timeliner
Timeliner is a Go CLI tool that archives your digital life—photos, social media posts, messages, and location history from services like Go…
103558abandoned
mozilla-services/heka
Heka is a Go-based tool for collecting data from many sources, performing in-flight processing, and delivering results to multiple destinat…
103394abandoned
harismuneer/Ultimate-Social-Scrapers
A collection of Python-based scraping tools that extract public data from Facebook, Instagram, and Twitter (X), including posts, media, fol…
443151abandoned
airingursb/bilibili-user
A Python web crawler that scrapes Bilibili user profiles (id, nickname, gender, avatar, level, birthday, location, etc.) and stores them in…
323090abandoned
WeRead2Notion
A Python tool that syncs WeChat Reading (WeRead) highlights and notes into a Notion database, typically run on a daily schedule via GitHub …
612919abandoned
jprante/elasticsearch-jdbc
A JDBC importer tool that fetches tabular data from relational databases via SQL queries and indexes it into Elasticsearch. It runs as a Ja…
102809abandoned
FeatureBaseDB/featurebase
FeatureBase (formerly Pilosa) is a distributed analytical database built entirely on bitmap indexes, offering low-latency SQL queries over …
102521abandoned
instill-ai/instill-core
Instill Core is a full-stack AI infrastructure tool for data, model, and pipeline orchestration, aimed at building AI-first applications wi…
762321abandoned
chiphuyen/lazynlp
A Python library for crawling, cleaning, and deduplicating web pages to build massive monolingual text datasets, suitable for training lang…
232284abandoned
shyiko/mysql-binlog-connector-java
A Java library for reading MySQL binary logs, both from binlog files and by tapping into the live MySQL replication stream. It supports GTI…
232269abandoned
ByConity/ByConity
ByConity is an open-source cloud-native distributed SQL data warehouse built on the ClickHouse 21.8 codebase, developed by ByteDance/Volcan…
102239abandoned
dominictarr/event-stream
EventStream is a Node.js library for composing pipelines of event streams using functional operations like map, filter, split, and merge. I…
102175abandoned
minimaxir/facebook-page-post-scraper
A Python script collection that scrapes all posts, reactions, and comments from public Facebook Pages and open Groups via the Facebook Grap…
102135abandoned
twitter/summingbird
Summingbird is a Scala library for writing MapReduce aggregation programs that look like native collection transformations and run on distr…
102123abandoned
onyx-platform/onyx
Onyx is a masterless, fault-tolerant distributed computation system written in pure Clojure that supports both batch and stream processing …
102051abandoned
Tencent/TubeMQ
TubeMQ was Tencent's high-performance message queue, donated to the Apache Software Foundation in 2019 and renamed Apache InLong. This repo…
231996abandoned
torodb/stampede
ToroDB Stampede replicates data from a MongoDB replica set into relational tables in PostgreSQL, translating the document structure into no…
231757abandoned
ClimbsRocks/auto_ml
auto_ml is a Python library for automated machine learning that handles feature engineering, model selection, and hyperparameter optimizati…
231654abandoned
discoproject/disco
Disco is an open-source distributed MapReduce framework written in Erlang with a Python job API, originally developed at Nokia Research Cen…
321630abandoned
stripe-archive/mosql
MoSQL is a Ruby tool that streams the contents of a MongoDB cluster into PostgreSQL, using an oplog tailer to keep the SQL mirror continuou…
101616abandoned
mongodb/mongo-hadoop
A Java library that lets MongoDB (or BSON backup files) serve as an input source or output destination for Hadoop MapReduce jobs, with inte…
101552abandoned
fossasia/open-event-scraper
A Python tool that parses Google spreadsheets and converts them into Open Event JSON format, originally built for FOSSASIA 2016 event data.…
101517abandoned
qinxuye/cola
Cola is a high-level distributed crawling framework in Python for scraping pages and extracting structured data from websites. The same cra…
101499abandoned
compose/transporter
Transporter is a Go CLI tool that syncs and transforms data between persistence engines such as MongoDB, PostgreSQL, Elasticsearch, MySQL, …
101441abandoned
mesos/spark
This is the original UC Berkeley AMPLab repository for Apache Spark, a fast cluster computing system supporting Java, Scala, and Python. Th…
321418abandoned
DocNow/twarc
twarc is a command line tool and Python library for collecting and archiving Twitter JSON data via the Twitter v1.1 and v2 APIs. It was dev…
501394abandoned
Jefferson-Henrique/GetOldTweets-python
A Python library that retrieves old tweets by mimicking the JSON calls Twitter Search makes in the browser, bypassing the official API's ti…
321340abandoned
nytlabs/streamtools
Streamtools is a graphical toolkit from The New York Times R&D Lab for exploring, analyzing, and modifying streams of real-time data. Users…
231312abandoned
yahoo/CaffeOnSpark
CaffeOnSpark is a Spark package that brings the Caffe deep learning framework to Hadoop and Spark clusters, enabling distributed neural net…
101261abandoned
uber-archive/AthenaX
AthenaX is a SQL-based streaming analytics platform open sourced by Uber, built on Apache Flink and Apache Calcite. It lets users run produ…
101223abandoned
foolcage/fooltrader
fooltrader is a Python quantitative analysis and trading framework that crawls, cleans, and structures market data (stocks, futures, forex,…
231198abandoned
killrweather/killrweather
KillrWeather is a Scala reference application demonstrating integration of Apache Spark Streaming, Apache Kafka, Apache Cassandra, and Akka…
321178abandoned
richardwilly98/elasticsearch-river-mongodb
An Elasticsearch plugin (river) that indexes MongoDB collections into Elasticsearch by tailing the MongoDB oplog of a replica set. It suppo…
321121abandoned
nfldb
A Python library (nflgame) providing an API to retrieve and read NFL Game Center JSON data, including real-time feeds useful for fantasy fo…
231082abandoned
jorgecarleitao/arrow2
A Rust implementation of the Apache Arrow columnar in-memory format, supporting IO with Parquet, Avro, CSV, IPC, and Flight, plus compute k…
101064abandoned
LockerProject/Locker
Locker is an open-source personal data platform ('the me platform') that collects a user's data from various online services, apps, and dev…
321063abandoned
littlstar/s3-lambda
A Node.js library that lets you run lambda-style functions (forEach, map, reduce, filter) over S3 objects with concurrency control. It enab…
321059abandoned
paulyoder/LinqToExcel
A .NET library that lets developers query Excel spreadsheets and CSV files using LINQ syntax, mapping worksheet rows to strongly-typed obje…
321058abandoned
klbostee/dumbo
Dumbo is a Python module that makes writing and running Hadoop Streaming programs easy, providing a convenient Python API for MapReduce pro…
321030abandoned
bcbio/bcbio-nextgen
bcbio-nextgen is a validated, community-developed pipeline toolkit for high-throughput sequencing analysis, including variant calling, RNA-…
231030abandoned
facebookarchive/bistro
Bistro is a C++ distributed task scheduler framework from Facebook that schedules and runs distributed tasks, including data-parallel jobs,…
101026abandoned
n8n
n8n is a fair-code workflow automation platform with native AI capabilities, combining a visual canvas with custom JavaScript/Python code. …
95202527active
infiniflow/ragflow
RAGFlow is an open-source Retrieval-Augmented Generation (RAG) engine that combines deep document understanding with agent orchestration to…
9389328active
unclecode/crawl4ai
Crawl4AI is an open-source Python library that crawls websites with a headless browser and converts pages into clean, LLM-ready Markdown fo…
8979471active
scrapy/scrapy
Scrapy is a fast, high-level web crawling and scraping framework for Python used to extract structured data from websites. It provides a fu…
9964048stable
sansan0/TrendRadar
TrendRadar is a self-hostable AI-powered public opinion and trending-news monitor that aggregates hot topics across multiple platforms and …
6161855active
LlamaIndex
LlamaIndex is a Python (with a TypeScript variant) framework for building LLM-powered applications, centered on ingesting, indexing, and qu…
9551883stable
microsoft/qlib
Qlib is an AI-oriented quantitative investment platform from Microsoft that supports the full quant research workflow, from data processing…
6647960active
NaiboWang/EasySpider
EasySpider is a free, open-source visual no-code web crawler and browser automation (RPA) tool where users design scraping tasks by clickin…
8744442active
ray-project/ray
Ray is a unified open-source framework for scaling AI and Python applications, consisting of a core distributed runtime (tasks, actors, obj…
9943614stable
DuckDB
DuckDB is an in-process analytical SQL database management system (OLAP) written in C++, designed to be fast, portable, and easy to embed. …
9740675stable
topoteretes/cognee
Cognee is an open-source Python library and platform that gives AI agents persistent long-term memory by ingesting data in any format and b…
9130281active
PrefectHQ/prefect
Prefect is an open-source workflow orchestration framework that turns Python functions into production-grade data pipelines using decorator…
9523693stable
recommenders-team/recommenders
A Python library and collection of Jupyter notebooks with best practices for building, evaluating, and operationalizing recommendation syst…
6721864active
openobserve/openobserve
OpenObserve is an open-source, cloud-native observability platform that unifies logs, metrics, traces, RUM, and LLM observability in a sing…
9321490active
xming521/WeClone
WeClone is an end-to-end Python framework for creating a personal AI digital twin by fine-tuning large language models on your exported cha…
8018171active
windmill-labs/windmill
Windmill is an open-source, self-hostable developer platform that turns scripts in Python, TypeScript, Go, Bash, SQL and other languages in…
9517687active
argoproj/argo-workflows
Argo Workflows is an open-source, container-native workflow engine for orchestrating parallel jobs on Kubernetes, implemented as a Kubernet…
9916939stable
apache/doris
Apache Doris is an open-source MPP-based real-time analytical database that delivers sub-second queries over massive datasets, combining a …
9915817stable
open-metadata/OpenMetadata
OpenMetadata is an open-source metadata management platform that unifies data cataloging, lineage, data quality, governance, and business s…
9514986active
microsoft/RD-Agent
RD-Agent is a Microsoft open-source framework that uses LLM-powered agents to automate research and development processes such as factor mi…
7514342active
apache/druid
Apache Druid is a high-performance, distributed, real-time analytics database written in Java for fast OLAP-style slice-and-dice queries on…
8914045stable
trinodb/trino
Trino is a fast, distributed ANSI SQL query engine for big data analytics, formerly known as PrestoSQL. It queries data in place across div…
9313183stable
StarRocks/starrocks
StarRocks is a high-performance distributed OLAP database and query engine built on an MPP architecture with a fully vectorized execution e…
9512043active
dataelement/bisheng
BISHENG is an open-source LLM DevOps (LLMOps) platform for building enterprise AI applications, offering GenAI workflow orchestration, RAG,…
9411911active
code4craft/webmagic
WebMagic is a scalable web crawler framework for Java covering the full crawl lifecycle: downloading, URL management, content extraction (X…
6011678active
cocoindex-io/cocoindex
CocoIndex is an open-source incremental data framework for AI, with a Rust engine and a declarative Python API that keeps sources (files, S…
8211409active
rerun-io/rerun
Rerun is an open-source SDK and viewer for logging, storing, querying, and visualizing multi-rate multimodal data such as images, point clo…
9911362active
semantica-agi/semantica
Semantica is a Python library providing graph-native infrastructure for building context graphs, knowledge graphs, and decision intelligenc…
8510920active
Netflix/metaflow
Metaflow is a human-centric Python framework from Netflix for building, managing, and deploying real-life AI/ML and data science systems. I…
9510245stable
databendlabs/databend
Databend is an open-source, cloud-native data warehouse built in Rust that runs entirely on object storage (S3, Azure, GCS). It unifies BI …
939423active
spring-projects/spring-ai
Spring AI is an application framework for AI engineering that provides Spring-friendly, portable abstractions for integrating AI models int…
959358active
Deep Lake
Deep Lake is an open-source database for AI that stores multimodal data (images, video, audio, text, embeddings, annotations) in a format o…
779228active
apache/iceberg
Apache Iceberg is a high-performance open table format for huge analytic datasets, bringing SQL table reliability to big data lakes. This r…
909177stable
GoogleCloudPlatform/knowledge-catalog
A collection of tools, agents, and samples for Google Cloud's Knowledge Catalog (formerly Dataplex), an AI-powered data catalog and metadat…
588913active
fluent/fluent-bit
Fluent Bit is a fast, lightweight telemetry agent written in C that collects, processes, and forwards logs, metrics, and traces from any so…
998062stable
pythonstock/stock
A full-stack stock analysis system built in Python using pandas, akshare, bokeh, stockstats, ta-lib, and tornado, deployed via Docker and d…
597864active
adithya-s-k/omniparse
OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into…
497815active
andeya/pholcus
Pholcus is a distributed, high-concurrency web crawler framework written in pure Go. It supports standalone, server, and client modes with …
907577active
dbgate/dbgate
DbGate is a cross-platform, open-source (no)SQL database manager for MySQL, PostgreSQL, SQL Server, MongoDB, SQLite, Redis and many other d…
947285active
flyteorg/flyte
Flyte is an open-source workflow orchestration platform for coordinating data, ML models, and AI agents at scale, authored in pure Python a…
957245active
Alluxio/alluxio
Alluxio is an open-source distributed caching and data orchestration platform that sits between compute frameworks (Spark, Presto, Trino, P…
317231stable
Zipstack/unstract
Unstract is an open-source, LLM-driven platform that extracts structured JSON data from unstructured documents such as PDFs, images, and sc…
887172active
rocketride-org/rocketride-server
RocketRide is an open-source AI pipeline engine with a high-throughput C++ runtime and 100+ Python-extensible nodes for building, debugging…
807090active
deepseek-ai/DeepSpec
DeepSpec is a full-stack Python codebase from DeepSeek for training and evaluating draft models used in speculative decoding of large langu…
547041active
Hazelcast
Hazelcast is a unified real-time data platform combining distributed stream processing with a fast, in-memory data store. It lets applicati…
826604stable
apache/camel
Apache Camel is an open-source integration framework offering 350+ connectors for databases, APIs, message brokers, and cloud services, wit…
776300stable

← prev page 5 / 7 next →