function: etl
671 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| aeon-toolkit/aeon aeon is a scikit-learn compatible Python toolkit for machine learning on time series, covering classification, regression, clustering, fore… | 91 | 1438 | active |
| movingpandas/movingpandas MovingPandas is a Python library for movement data exploration and analysis, providing trajectory data structures built on Pandas, GeoPanda… | 98 | 1408 | active |
| apache/iceberg-rust A Rust implementation of the Apache Iceberg open table format for managing large analytic datasets. It provides crates for the core Iceberg… | 90 | 1390 | active |
| stellar/go-stellar-sdk The official Go SDK for the Stellar blockchain network, maintained by the Stellar Development Foundation. It provides transaction building,… | 99 | 1383 | active |
| CIRCL/AIL-framework AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T… | 67 | 1378 | active |
| quiltdata/quilt Quilt is a scientific data management platform built on AWS that turns S3 data into deeply versioned, metadata-rich data packages for teams… | 67 | 1370 | active |
| loggie-io/loggie Loggie is a lightweight, high-performance, cloud-native log collection agent and aggregator written in Go. It uses Kubernetes CRDs to confi… | 23 | 1333 | active |
| StarlightSearch/EmbedAnything EmbedAnything is a high-performance, memory-safe embedding pipeline written in Rust (with Python bindings) that generates embeddings from t… | 88 | 1305 | active |
| logicalclocks/hopsworks Hopsworks is an open-source, data-intensive AI platform (an 'AI Lakehouse') built around a Python-centric Feature Store with online/offline… | 26 | 1303 | active |
| pyexcel/pyexcel pyexcel is a Python library providing a single unified API for reading, manipulating, and writing tabular data across spreadsheet formats i… | 84 | 1291 | active |
| wq/django-rest-pandas Django REST Pandas is a Python library that serves pandas DataFrames through the Django REST Framework as a model-driven data API. It rende… | 52 | 1277 | stable |
| water8394/flink-recommandSystem-demo A real-time product recommendation system built on Apache Flink, demonstrating streaming computation of product popularity, user profiles, … | 32 | 4480 | maintenance |
| tatuylonen/wiktextract A Python package and CLI tool that parses Wiktionary XML dump files and extracts structured dictionary data (glosses, translations, pronunc… | 76 | 1251 | active |
| DefiLlama/DefiLlama-Adapters A collection of community-contributed TVL (Total Value Locked) adapters for DefiLlama, written in JavaScript, that compute on-chain protoco… | 72 | 1250 | active |
| nfstream/nfstream NFStream is a multiplatform Python framework for fast, flexible network flow data analysis from live interfaces or pcap files. It provides … | 81 | 1218 | stable |
| ardha27/AI-Song-Cover-RVC A collection of Google Colab and Kaggle notebooks that form an all-in-one toolkit for creating AI song covers with RVC (Retrieval-based Voi… | 68 | 1207 | active |
| opensemanticsearch/open-semantic-search An open-source integrated search server and ETL framework for processing, analyzing, and exploring large document collections. It combines … | 40 | 1203 | active |
| LifeArchiveProject/BilibiliHistoryFetcher A Python/FastAPI backend tool that fetches, stores, and analyzes a user's Bilibili watch history, favorites, dynamics, comments, and intera… | 81 | 1201 | active |
| Cerebras/modelzoo Cerebras Model Zoo is a collection of reference deep learning model implementations (Llama, Mixtral, DINOv2, Llava, etc.) with configs and … | 77 | 1193 | active |
| mdbtools/mdbtools MDB Tools is a set of C libraries and command-line utilities for reading Microsoft Access (.mdb/.accdb) database files on Unix-like systems… | 51 | 1168 | active |
| vincentlaucsb/csv-parser A modern C++17 CSV parser and serializer that is RFC 4180 compliant and balances ease of use with high performance. It supports streaming a… | 95 | 1121 | active |
| datadreamer-dev/DataDreamer DataDreamer is an open-source Python library for prompting LLMs, generating synthetic datasets, and training or aligning models in reproduc… | 33 | 1117 | active |
| elixir-crawly/crawly Crawly is a high-level web crawling and scraping framework for Elixir, modeled after Scrapy, where developers define spiders that fetch pag… | 37 | 1114 | active |
| rstudio/pointblank pointblank is an R package for methodically validating data in data frames and database tables, producing data quality reports. It also mai… | 86 | 1047 | active |
| D2I-CUHKSZ/MicroWorld MicroWorld is a lightweight Python engine that turns multi-modal event materials (documents, images, videos, graph signals) into structured… | 53 | 1039 | active |
| NVIDIA-NeMo/Skills Nemo-Skills is a collection of Python pipelines for improving the skills of large language models, covering synthetic data generation, mode… | 70 | 1031 | active |
| GlareDB/glaredb GlareDB is a lightweight, fast SQL database engine built in Rust for running analytics queries directly against files and object storage li… | 57 | 1021 | active |
| towhee-io/towhee Towhee is a Python framework for building ETL pipelines that process unstructured data (images, video, text, audio) into embeddings using s… | 23 | 3454 | maintenance |
| facebookresearch/PyTorch-BigGraph PyTorch-BigGraph is a distributed system for learning embeddings of very large graph-structured data, scaling to billions of entities and t… | 10 | 3454 | maintenance |
| Logflare/logflare Logflare is a centralized structured log ingestion and querying service that streams log events into BigQuery (or its own backend) with aut… | 96 | 1004 | active |
| airbnb/streamalert StreamAlert is a serverless, real-time data analysis framework from Airbnb for ingesting, analyzing, and alerting on log data from any envi… | 23 | 2889 | maintenance |
| Teevity/ice Ice is a self-hosted web application (Grails-based, Java) that processes AWS detailed billing files and renders interactive usage and cost … | 23 | 2878 | maintenance |
| brianway/webporter webporter is a Java crawler application built on the webmagic framework that demonstrates a complete pipeline of data crawling, persistence… | 23 | 2766 | maintenance |
| pipelinedb/pipelinedb PipelineDB is a PostgreSQL extension for high-performance time-series aggregation, letting you define continuous SQL queries that increment… | 23 | 2662 | maintenance |
| geekyouth/SZT-bigdata A big data passenger flow analysis system for the Shenzhen Metro, built on Shenzhen Tong smart-card swipe data. It demonstrates ETL and ana… | 60 | 2475 | maintenance |
| influxdata/kapacitor Kapacitor is an open-source framework for processing, monitoring, and alerting on time series data, part of the InfluxData TICK stack. It u… | 90 | 2375 | maintenance |
| dflemstr/rq rq (Record Query) is a Rust command-line tool for ad-hoc analysis and transformation of streams of structured records, similar to awk/sed b… | 23 | 2299 | maintenance |
| sfu-db/dataprep DataPrep is a Python library for low-code data preparation, offering modules to collect data from common APIs (connector), run fast explora… | 23 | 2248 | maintenance |
| epfLLM/meditron Meditron is a suite of open-source medical large language models (7B and 70B) adapted from Llama-2 via continued pretraining on a curated m… | 27 | 2208 | maintenance |
| DTStack/flinkStreamSQL FlinkStreamSQL is a Java framework built on Apache Flink that extends Flink's real-time SQL with custom create table/view/function syntax a… | 23 | 2051 | maintenance |
| Qihoo360/poseidon Poseidon is a distributed log search platform from Qihoo 360 that builds inverted indexes over Hadoop/HDFS-stored logs and serves sub-secon… | 32 | 1981 | maintenance |
| HazyResearch/deepdive DeepDive is a Stanford-developed system for extracting structured data from unstructured sources and building knowledge bases using distant… | 23 | 1979 | maintenance |
| ICT-BDA/EasyML EasyML is a general-purpose dataflow-based machine learning platform where tasks are defined as directed acyclic graphs of operations. It i… | 23 | 1976 | maintenance |
| minimaxir/automl-gs automl-gs is a Python AutoML tool that takes an input CSV and a target prediction field and automatically generates a trained machine learn… | 23 | 1866 | maintenance |
| ScottfreeLLC/AlphaPy AlphaPy is a Python machine learning framework built on scikit-learn, pandas, Keras, XGBoost, LightGBM, and CatBoost for building classific… | 40 | 1745 | maintenance |
| re-data/re-data re_data is an open-source data reliability framework built as a dbt package for the modern data stack. It computes data quality metrics, de… | 23 | 1569 | maintenance |
| fossasia/event-collect A Python CLI tool that scrapes event website listings (e.g., EventBrite search results) and converts them into the Open Event JSON format. … | 10 | 1505 | maintenance |
| apache/carbondata Apache CarbonData is an indexed columnar data file format and store for fast analytics on big data platforms like Apache Hadoop and Apache … | 75 | 1452 | maintenance |
| liukelin/canal_mysql_nosql_sync A demo application showing real-time synchronization of MySQL data into Redis, Memcached, and MongoDB using Alibaba's Canal binlog parser a… | 23 | 1410 | maintenance |
| istresearch/scrapy-cluster Scrapy Cluster is a distributed web scraping framework built on Scrapy that uses Redis to coordinate crawl requests and Kafka as a data bus… | 10 | 1225 | maintenance |
| astroML/astroML AstroML is a Python library for machine learning, statistics, and data mining aimed at astronomy and astrophysics, built on numpy, scipy, s… | 32 | 1200 | maintenance |
| business-science/ai-data-science-team A Python library of specialized LLM-powered agents for common data science workflows such as data loading, cleaning, wrangling, visualizati… | 60 | 5385 | experimental |
| Boerderij/Varken Varken is a standalone Python application that aggregates data from Plex ecosystem tools (Sonarr, Radarr, Tautulli, Lidarr, Ombi, SickChill… | 23 | 1176 | maintenance |
| MarcosMeli/FileHelpers FileHelpers is a free, MIT-licensed .NET library for reading and writing strongly typed records from fixed-length or delimited flat files, … | 23 | 1153 | maintenance |
| acikyazilimagi/afet-org An open-source earthquake relief platform (depremyardim.com / afetharita.com) that aggregates calls for help from Twitter, WhatsApp, Telegr… | 31 | 1079 | maintenance |
| dfm/osrc The Open Source Report Card (OSRC) is a web application that generates reports about a user's open-source activity, originally based on Git… | 10 | 1030 | maintenance |
| tinyfish-io/bigset-oss BigSet is a self-hostable application that turns a natural-language sentence into a structured, regularly refreshed dataset by dispatching … | 76 | 1684 | experimental |
| twintproject/twint Twint is a Python CLI tool and library that scrapes tweets, followers, following, and likes from Twitter without using the official API or … | 10 | 16398 | abandoned |
| BurntSushi/xsv xsv is a fast, composable command-line toolkit written in Rust for indexing, slicing, analyzing, splitting, and joining CSV files. It offer… | 10 | 10757 | abandoned |
| axa-group/Parsr Parsr is a document parsing and extraction toolchain that transforms PDFs, images, docx, and eml files into clean, enriched structured data… | 56 | 6178 | abandoned |
| wuhan2020/wuhan2020 A volunteer-built information collection platform for the 2020 Wuhan COVID-19 outbreak, aggregating data on hospitals, hotels, logistics, d… | 32 | 5913 | abandoned |
| timescale/pgai pgai is a Python library and set of PostgreSQL tools from Timescale that turns PostgreSQL into a retrieval engine for RAG and agentic appli… | 10 | 5806 | abandoned |
| truefoundry/cognita Cognita is an open-source RAG (Retrieval Augmented Generation) framework by TrueFoundry for building modular, production-ready RAG applicat… | 10 | 4418 | abandoned |
| run-llama/llama_cloud_services LlamaCloud Services provides cloud-hosted document parsing and knowledge agent tooling, including the LlamaParse document parser that conve… | 71 | 4262 | abandoned |
| intel/BigDL BigDL is Intel's distributed deep learning library that scales TensorFlow, Keras, and PyTorch workloads on Apache Spark, Flink, and Ray, wi… | 10 | 2698 | abandoned |
| Azure-Samples/graphrag-accelerator A one-click Azure deployment accelerator that hosts the GraphRAG knowledge-graph-powered RAG system as a scalable API service. It builds on… | 10 | 2409 | abandoned |
| dora-team/fourkeys Four Keys is a self-hostable platform from Google's DORA team that collects events from development environments like GitHub and GitLab via… | 10 | 2236 | abandoned |
| microsoft/kernel-memory Kernel Memory is a multi-modal AI service and .NET library for indexing datasets through hybrid data pipelines, enabling RAG, semantic sear… | 69 | 2235 | abandoned |
| instacart/lore Lore is a Python machine learning framework from Instacart designed to make ML approachable for software engineers and maintainable for ML … | 23 | 1543 | abandoned |
| alibaba/mdrill Mdrill is an open-source distributed OLAP (online analytical processing) engine from Alibaba's AdMom team, built in Java on top of JStorm/H… | 10 | 1543 | abandoned |
| benitoro/stockholm A Python framework that crawls Shanghai and Shenzhen A-share stock market data from Yahoo YQL and Sina Finance, and tests stock-picking str… | 32 | 1517 | abandoned |
← prev page 7 / 7