Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: etl

671 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
aeon-toolkit/aeon
aeon is a scikit-learn compatible Python toolkit for machine learning on time series, covering classification, regression, clustering, fore…
911438active
movingpandas/movingpandas
MovingPandas is a Python library for movement data exploration and analysis, providing trajectory data structures built on Pandas, GeoPanda…
981408active
apache/iceberg-rust
A Rust implementation of the Apache Iceberg open table format for managing large analytic datasets. It provides crates for the core Iceberg…
901390active
stellar/go-stellar-sdk
The official Go SDK for the Stellar blockchain network, maintained by the Stellar Development Foundation. It provides transaction building,…
991383active
CIRCL/AIL-framework
AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T…
671378active
quiltdata/quilt
Quilt is a scientific data management platform built on AWS that turns S3 data into deeply versioned, metadata-rich data packages for teams…
671370active
loggie-io/loggie
Loggie is a lightweight, high-performance, cloud-native log collection agent and aggregator written in Go. It uses Kubernetes CRDs to confi…
231333active
StarlightSearch/EmbedAnything
EmbedAnything is a high-performance, memory-safe embedding pipeline written in Rust (with Python bindings) that generates embeddings from t…
881305active
logicalclocks/hopsworks
Hopsworks is an open-source, data-intensive AI platform (an 'AI Lakehouse') built around a Python-centric Feature Store with online/offline…
261303active
pyexcel/pyexcel
pyexcel is a Python library providing a single unified API for reading, manipulating, and writing tabular data across spreadsheet formats i…
841291active
wq/django-rest-pandas
Django REST Pandas is a Python library that serves pandas DataFrames through the Django REST Framework as a model-driven data API. It rende…
521277stable
water8394/flink-recommandSystem-demo
A real-time product recommendation system built on Apache Flink, demonstrating streaming computation of product popularity, user profiles, …
324480maintenance
tatuylonen/wiktextract
A Python package and CLI tool that parses Wiktionary XML dump files and extracts structured dictionary data (glosses, translations, pronunc…
761251active
DefiLlama/DefiLlama-Adapters
A collection of community-contributed TVL (Total Value Locked) adapters for DefiLlama, written in JavaScript, that compute on-chain protoco…
721250active
nfstream/nfstream
NFStream is a multiplatform Python framework for fast, flexible network flow data analysis from live interfaces or pcap files. It provides …
811218stable
ardha27/AI-Song-Cover-RVC
A collection of Google Colab and Kaggle notebooks that form an all-in-one toolkit for creating AI song covers with RVC (Retrieval-based Voi…
681207active
opensemanticsearch/open-semantic-search
An open-source integrated search server and ETL framework for processing, analyzing, and exploring large document collections. It combines …
401203active
LifeArchiveProject/BilibiliHistoryFetcher
A Python/FastAPI backend tool that fetches, stores, and analyzes a user's Bilibili watch history, favorites, dynamics, comments, and intera…
811201active
Cerebras/modelzoo
Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and …
771193active
mdbtools/mdbtools
MDB Tools is a set of C libraries and command-line utilities for reading Microsoft Access (.mdb/.accdb) database files on Unix-like systems…
511168active
vincentlaucsb/csv-parser
A modern C++17 CSV parser and serializer that is RFC 4180 compliant and balances ease of use with high performance. It supports streaming a…
951121active
datadreamer-dev/DataDreamer
DataDreamer is an open-source Python library for prompting LLMs, generating synthetic datasets, and training or aligning models in reproduc…
331117active
elixir-crawly/crawly
Crawly is a high-level web crawling and scraping framework for Elixir, modeled after Scrapy, where developers define spiders that fetch pag…
371114active
rstudio/pointblank
pointblank is an R package for methodically validating data in data frames and database tables, producing data quality reports. It also mai…
861047active
D2I-CUHKSZ/MicroWorld
MicroWorld is a lightweight Python engine that turns multi-modal event materials (documents, images, videos, graph signals) into structured…
531039active
NVIDIA-NeMo/Skills
Nemo-Skills is a collection of Python pipelines for improving the skills of large language models, covering synthetic data generation, mode…
701031active
GlareDB/glaredb
GlareDB is a lightweight, fast SQL database engine built in Rust for running analytics queries directly against files and object storage li…
571021active
towhee-io/towhee
Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s…
233454maintenance
facebookresearch/PyTorch-BigGraph
PyTorch-BigGraph is a distributed system for learning embeddings of very large graph-structured data, scaling to billions of entities and t…
103454maintenance
Logflare/logflare
Logflare is a centralized structured log ingestion and querying service that streams log events into BigQuery (or its own backend) with aut…
961004active
airbnb/streamalert
StreamAlert is a serverless, real-time data analysis framework from Airbnb for ingesting, analyzing, and alerting on log data from any envi…
232889maintenance
Teevity/ice
Ice is a self-hosted web application (Grails-based, Java) that processes AWS detailed billing files and renders interactive usage and cost …
232878maintenance
brianway/webporter
webporter is a Java crawler application built on the webmagic framework that demonstrates a complete pipeline of data crawling, persistence…
232766maintenance
pipelinedb/pipelinedb
PipelineDB is a PostgreSQL extension for high-performance time-series aggregation, letting you define continuous SQL queries that increment…
232662maintenance
geekyouth/SZT-bigdata
A big data passenger flow analysis system for the Shenzhen Metro, built on Shenzhen Tong smart-card swipe data. It demonstrates ETL and ana…
602475maintenance
influxdata/kapacitor
Kapacitor is an open-source framework for processing, monitoring, and alerting on time series data, part of the InfluxData TICK stack. It u…
902375maintenance
dflemstr/rq
rq (Record Query) is a Rust command-line tool for ad-hoc analysis and transformation of streams of structured records, similar to awk/sed b…
232299maintenance
sfu-db/dataprep
DataPrep is a Python library for low-code data preparation, offering modules to collect data from common APIs (connector), run fast explora…
232248maintenance
epfLLM/meditron
Meditron is a suite of open-source medical large language models (7B and 70B) adapted from Llama-2 via continued pretraining on a curated m…
272208maintenance
DTStack/flinkStreamSQL
FlinkStreamSQL is a Java framework built on Apache Flink that extends Flink's real-time SQL with custom create table/view/function syntax a…
232051maintenance
Qihoo360/poseidon
Poseidon is a distributed log search platform from Qihoo 360 that builds inverted indexes over Hadoop/HDFS-stored logs and serves sub-secon…
321981maintenance
HazyResearch/deepdive
DeepDive is a Stanford-developed system for extracting structured data from unstructured sources and building knowledge bases using distant…
231979maintenance
ICT-BDA/EasyML
EasyML is a general-purpose dataflow-based machine learning platform where tasks are defined as directed acyclic graphs of operations. It i…
231976maintenance
minimaxir/automl-gs
automl-gs is a Python AutoML tool that takes an input CSV and a target prediction field and automatically generates a trained machine learn…
231866maintenance
ScottfreeLLC/AlphaPy
AlphaPy is a Python machine learning framework built on scikit-learn, pandas, Keras, XGBoost, LightGBM, and CatBoost for building classific…
401745maintenance
re-data/re-data
re_data is an open-source data reliability framework built as a dbt package for the modern data stack. It computes data quality metrics, de…
231569maintenance
fossasia/event-collect
A Python CLI tool that scrapes event website listings (e.g., EventBrite search results) and converts them into the Open Event JSON format. …
101505maintenance
apache/carbondata
Apache CarbonData is an indexed columnar data file format and store for fast analytics on big data platforms like Apache Hadoop and Apache …
751452maintenance
liukelin/canal_mysql_nosql_sync
A demo application showing real-time synchronization of MySQL data into Redis, Memcached, and MongoDB using Alibaba's Canal binlog parser a…
231410maintenance
istresearch/scrapy-cluster
Scrapy Cluster is a distributed web scraping framework built on Scrapy that uses Redis to coordinate crawl requests and Kafka as a data bus…
101225maintenance
astroML/astroML
AstroML is a Python library for machine learning, statistics, and data mining aimed at astronomy and astrophysics, built on numpy, scipy, s…
321200maintenance
business-science/ai-data-science-team
A Python library of specialized LLM-powered agents for common data science workflows such as data loading, cleaning, wrangling, visualizati…
605385experimental
Boerderij/Varken
Varken is a standalone Python application that aggregates data from Plex ecosystem tools (Sonarr, Radarr, Tautulli, Lidarr, Ombi, SickChill…
231176maintenance
MarcosMeli/FileHelpers
FileHelpers is a free, MIT-licensed .NET library for reading and writing strongly typed records from fixed-length or delimited flat files, …
231153maintenance
acikyazilimagi/afet-org
An open-source earthquake relief platform (depremyardim.com / afetharita.com) that aggregates calls for help from Twitter, WhatsApp, Telegr…
311079maintenance
dfm/osrc
The Open Source Report Card (OSRC) is a web application that generates reports about a user's open-source activity, originally based on Git…
101030maintenance
tinyfish-io/bigset-oss
BigSet is a self-hostable application that turns a natural-language sentence into a structured, regularly refreshed dataset by dispatching …
761684experimental
twintproject/twint
Twint is a Python CLI tool and library that scrapes tweets, followers, following, and likes from Twitter without using the official API or …
1016398abandoned
BurntSushi/xsv
xsv is a fast, composable command-line toolkit written in Rust for indexing, slicing, analyzing, splitting, and joining CSV files. It offer…
1010757abandoned
axa-group/Parsr
Parsr is a document parsing and extraction toolchain that transforms PDFs, images, docx, and eml files into clean, enriched structured data…
566178abandoned
wuhan2020/wuhan2020
A volunteer-built information collection platform for the 2020 Wuhan COVID-19 outbreak, aggregating data on hospitals, hotels, logistics, d…
325913abandoned
timescale/pgai
pgai is a Python library and set of PostgreSQL tools from Timescale that turns PostgreSQL into a retrieval engine for RAG and agentic appli…
105806abandoned
truefoundry/cognita
Cognita is an open-source RAG (Retrieval Augmented Generation) framework by TrueFoundry for building modular, production-ready RAG applicat…
104418abandoned
run-llama/llama_cloud_services
LlamaCloud Services provides cloud-hosted document parsing and knowledge agent tooling, including the LlamaParse document parser that conve…
714262abandoned
intel/BigDL
BigDL is Intel's distributed deep learning library that scales TensorFlow, Keras, and PyTorch workloads on Apache Spark, Flink, and Ray, wi…
102698abandoned
Azure-Samples/graphrag-accelerator
A one-click Azure deployment accelerator that hosts the GraphRAG knowledge-graph-powered RAG system as a scalable API service. It builds on…
102409abandoned
dora-team/fourkeys
Four Keys is a self-hostable platform from Google's DORA team that collects events from development environments like GitHub and GitLab via…
102236abandoned
microsoft/kernel-memory
Kernel Memory is a multi-modal AI service and .NET library for indexing datasets through hybrid data pipelines, enabling RAG, semantic sear…
692235abandoned
instacart/lore
Lore is a Python machine learning framework from Instacart designed to make ML approachable for software engineers and maintainable for ML …
231543abandoned
alibaba/mdrill
Mdrill is an open-source distributed OLAP (online analytical processing) engine from Alibaba's AdMom team, built in Java on top of JStorm/H…
101543abandoned
benitoro/stockholm
A Python framework that crawls Shanghai and Shenzhen A-share stock market data from Yahoo YQL and Sina Finance, and tests stock-picking str…
321517abandoned

← prev page 7 / 7