Ross ROSS = Recommend OSS · open-source software intelligence for agents

resource: etl

121 resources, primary matches first, then adoption-weighted; health v2 shown.

ResourceHealth v2StarsMaturity
DataTalksClub/data-engineering-zoomcamp
A free 9-week open-source course teaching data engineering fundamentals by building an end-to-end production-ready data pipeline. It includ…
7745005active
andkret/Cookbook
The Data Engineering Cookbook is a free, open-source book by Andreas Kretz covering the skills, tools, and best practices needed to become …
7515371active
zhisheng17/flink-learning
A comprehensive Apache Flink learning repository with example code covering DataStream API, Table API & SQL, connectors, metrics, and produ…
6915095active
DataTalksClub/machine-learning-zoomcamp
A free 4-month cohort-based course from DataTalks.Club teaching machine learning engineering, covering the full path from problem framing a…
6814031active
joevess/IPTV
An automatically updated collection of IPTV live-stream playlists (m3u8) aggregating sources from haoqu, TVBox, and other public sources, s…
2810276active
igorbarinov/awesome-data-engineering
A curated awesome-list of data engineering tools, databases, frameworks, and resources for software developers. It organizes links across c…
748992active
rfordatascience/tidytuesday
TidyTuesday is a weekly social data project that publishes real-world datasets for people to practice data wrangling, visualization, and mo…
678359active
pditommaso/awesome-pipeline
A curated awesome-list cataloging pipeline toolkits, workflow engines, and data orchestration frameworks. It serves as a discovery resource…
756624active
togethercomputer/RedPajama-Data
RedPajama-Data provides code and pipelines for building RedPajama-V2, an open dataset with over 30 trillion tokens of web text for training…
684980active
mlabonne/llm-datasets
A curated list of datasets and tools for post-training large language models, covering supervised fine-tuning, preference, and reasoning da…
604757active
tensorflow/datasets
TensorFlow Datasets (TFDS) is a curated collection of ready-to-use public datasets exposed as tf.data.Dataset objects for TensorFlow, JAX, …
854581stable
decodingai-magazine/llm-twin-course
A free hands-on course (source code plus 12 lessons) that teaches how to build an end-to-end production-ready LLM and RAG system by creatin…
604385active
electricitymaps/electricitymaps-contrib
A community-maintained collection of Python parsers that collect and standardize electricity data (production, exchanges, prices) from offi…
894025active
jghoman/awesome-apache-airflow
A curated awesome-list of resources about Apache Airflow, the workflow orchestration platform. It collects tutorials, deployment solutions,…
693925active
Jon-Becker/prediction-market-analysis
A Python framework for collecting and analyzing prediction market data from Polymarket and Kalshi, including the largest publicly available…
603774active
xyflow/awesome-node-based-uis
A curated awesome-list of resources, libraries, and tools for building node-based UIs (node editors, visual programming environments, and w…
453664active
pawl/awesome-etl
A curated awesome-list of ETL (extract, transform, load) frameworks, libraries, and software, favoring mainstream, well-supported open-sour…
683585active
MIT-LCP/mimic-code
A community-maintained repository of code for building, loading, and analyzing the MIMIC family of critical care databases (MIMIC-III, MIMI…
793355active
openaddresses/openaddresses
OpenAddresses is a global, openly licensed collection of address, cadastral parcel, building footprint, and street centerline data sources.…
773242active
igrigorik/gharchive.org
GH Archive records the public GitHub timeline into hourly JSON event archives and makes them freely downloadable, plus available as a publi…
433072active
GoogleCloudPlatform/professional-services
A collection of example solutions, tools, and reference architectures developed by Google Cloud's Professional Services team, primarily in …
763065active
danielbeach/data-engineering-practice
A collection of hands-on data engineering practice problems covering Python data processing, file formats, SQL, Postgres, PySpark, and data…
322845active
The-Japan-DataScientist-Society/100knocks-preprocess
A collection of 100 structured data processing exercises (Data Science 100 Knocks) from the Japan Data Scientist Society, with practice pro…
602532active
binance/binance-public-data
A repository documenting and providing access to Binance's public historical market data, downloadable as daily or monthly files for spot, …
322462active
thenaturalist/awesome-business-intelligence
A curated awesome-list of business intelligence tools, split into SaaS and open-source categories covering analytics clients, data visualiz…
322314active
toddwschneider/nyc-taxi-data
A collection of R and shell scripts to download, process, and import 3+ billion NYC taxi and for-hire vehicle (Uber, Lyft) trip records int…
492081active
confluentinc/examples
A curated collection of demos and examples showcasing Apache Kafka, Apache Flink, and Confluent Platform event streaming, including Conflue…
772064active
awslabs/open-data-registry
A registry of publicly available datasets that can be accessed via AWS resources, with each dataset described in a YAML metadata file. It p…
771774active
scrollmapper/bible_databases
A collection of 140 Bible translations distributed as databases in multiple formats including MySQL, SQLite, CSV, JSON, YAML, TXT, Markdown…
731675active
oxylabs/scrape-google-python
A Python tutorial repository from Oxylabs demonstrating how to scrape Google search results (SERPs) using Oxylabs' SERP Scraper API. It inc…
631629active
AlexTheAnalyst/PortfolioProjects
A collection of data analyst portfolio projects with code and queries from Alex The Analyst's tutorials, primarily in Jupyter Notebooks. It…
321568active
aws-samples/aws-glue-samples
A collection of code samples, tutorials, and utilities demonstrating AWS Glue, Amazon's serverless data integration and ETL service. It cov…
761537active
Marvomatic/n8n-templates
A collection of free and premium n8n workflow templates for automating SEO, content optimization, and data analysis tasks. Templates integr…
441537active
datawhalechina/hands-on-data-analysis
An open-source Chinese-language hands-on course from Datawhale that teaches data analysis through project-based Jupyter notebooks covering …
321524active
duneanalytics/spellbook
Spellbook is a community-maintained collection of dbt SQL models that define curated, decoded datasets (like dex.trades and NFT data) on Du…
671513active
TurboWay/bigdata_analyse
A collection of hands-on big data analysis projects in Python, SQL, and HQL, each with documentation, datasets, and full pipelines from cle…
325325maintenance
GoogleCloudPlatform/data-science-on-gcp
Companion source code repository for the O'Reilly book 'Data Science on the Google Cloud Platform' by Valliappa Lakshmanan, provided as Jup…
631427active
chenditc/investment_data
A crowdsourced Chinese stock market dataset (A-shares) maintained in a Dolt database and exported to Microsoft Qlib binary format via Pytho…
951423active
lvgalvao/data-engineering-roadmap
The official repository of a professional training program in Data Engineering and AI (university extension) by Jornada de Dados. It contai…
631394active
spark-examples/pyspark-examples
A collection of runnable Python example scripts demonstrating PySpark RDD, DataFrame, and SQL operations, companion to the SparkByExamples …
571365active
singer-io/getting-started
The official getting-started guide and documentation for the Singer open-source ETL standard, which defines how extraction scripts (Taps) a…
481345active
Data-Learn/data-engineering
Data Learn is a free open educational resource and course series teaching data engineering and analytics, covering BI tools, databases, ETL…
401332active
datascale-ai/data_engineering_book
An open-source book (with GitHub Pages site and runnable Python code) on data engineering for large language models, covering pretraining d…
651300active
mahmoudparsian/pyspark-tutorial
A tutorial repository of Jupyter Notebook examples teaching basic distributed algorithms using PySpark, the Python API for Apache Spark. It…
431279active
MrSuiChuan/data-warehouse-learning
A Chinese-language educational project teaching real-time and offline data warehouse construction for an e-commerce system, with hands-on c…
611218active
codemayq/chinese-chatbot-corpus
A curated collection and unified processing pipeline for publicly available Chinese chit-chat conversation corpora, aggregating eight sourc…
324194maintenance
DataExpert-io/llm-driven-data-engineering
A public educational repository accompanying DataExpert's bootcamp course on LLM-driven data engineering concepts. It contains Python labs …
281156active
databricks/learning-spark
Example code accompanying the O'Reilly 'Learning Spark' book, with implementations in Java, Scala, and Python. The examples target Spark 1.…
733893maintenance
elastic/ecs
Elastic Common Schema (ECS) is an open-source specification defining a common set of fields, data types, and usage hierarchies for ingestin…
941119active
evansiroky/timezone-boundary-builder
A tool and dataset project that builds the world's timezone boundaries from OpenStreetMap data, released as shapefiles and GeoJSON. Each bo…
891114active
mahmoudparsian/data-algorithms-book
The official source code repository for Mahmoud Parsian's O'Reilly books 'Data Algorithms' and related titles, containing MapReduce, Spark,…
321082active
apache/flink-training
A collection of hands-on lab exercises and solutions accompanying the official Apache Flink training documentation. It provides a Gradle-ba…
691048active
OpenSourceAP/CrossSection
Code and data accompanying Chen and Zimmermann's paper on open source cross-sectional asset pricing. It reproduces dozens of stock-level pr…
491034active
tomwhite/hadoop-book
Example source code accompanying O'Reilly's 'Hadoop: The Definitive Guide' by Tom White, covering the Apache Hadoop ecosystem including Had…
323498maintenance
data-science-on-aws/data-science-on-aws
Companion Jupyter Notebook repository for the O'Reilly book 'Data Science on AWS', containing end-to-end AI/ML pipeline examples using Amaz…
323433maintenance
databricks/Spark-The-Definitive-Guide
The official code repository for the book 'Spark: The Definitive Guide' by Bill Chambers and Matei Zaharia, containing example code and dat…
323144maintenance
yyx990803/build-your-own-mint
A demo/tutorial project by Evan You showing how to build a personal finance analytics pipeline using Plaid bank data, Google Sheets, and Ci…
322538maintenance
qiurunze123/threadandjuc
A Java educational project teaching multithreading and concurrency from basics to advanced, culminating in a 'three-high' (high-availabilit…
322160maintenance
AlexIoannides/pyspark-example-project
A reference example project demonstrating best practices for structuring PySpark ETL jobs and applications. It shows how to organize code f…
102119maintenance
chris1610/pbpython
A collection of code, Jupyter notebooks, and examples accompanying the Practical Business Python blog (pbpython.com), which teaches Python …
321992maintenance
DXY-COVID-19
A time-series data warehouse of COVID-19 (2019-nCoV) infection statistics for China, scraped from Dingxiangyuan (DXY) and published as CSV/…
101971maintenance
san089/Udacity-Data-Engineering-Projects
A collection of Udacity Data Engineering Nanodegree projects covering data modeling with Postgres and Cassandra, cloud data warehousing wit…
321969maintenance
EleutherAI/the-pile
The Pile is a large, diverse, open-source language modeling dataset composed of many smaller text sources combined together. This repositor…
321673maintenance
jadianes/spark-py-notebooks
A collection of Jupyter/IPython notebooks teaching Apache Spark with Python (pySpark), from RDD basics to MLlib machine learning. It uses t…
321659maintenance
gururise/AlpacaDataCleaned
A cleaned and curated version of the Stanford Alpaca instruction-tuning dataset used to train the Alpaca LLM. It fixes quality issues in th…
621606maintenance
confluentinc/demo-scene
A collection of scripts, Docker Compose setups, and sample code supporting Confluent's demos, talks, and blogs about Apache Kafka and its e…
671565maintenance
sryza/aas
Source code examples accompanying the O'Reilly book 'Advanced Analytics with Spark', written in Scala and built with Maven. It provides run…
231525maintenance
griffithlab/rnaseq_tutorial
An educational web resource and working pipeline for RNA-seq analysis on the cloud, covering file formats, reference genomes, expression an…
231435maintenance
gtoonstra/etl-with-airflow
A community-maintained collection of ETL best practices, usage patterns, and examples for Apache Airflow, published as documentation with a…
321356maintenance
eleanorlutz/asteroids_atlas_of_space
An open-source tutorial and dataset for creating a detailed map of asteroid orbits in the solar system using NASA data. It includes Python …
321307maintenance
WillKoehrsen/machine-learning-project-walkthrough
A Jupyter Notebook-based walkthrough demonstrating a complete end-to-end machine learning solution on a real-world dataset. It shows how al…
321301maintenance
rsvp/fecon235
A collection of Jupyter notebooks for financial economics research, providing high-level Python interfaces to economic data sources like FR…
321275maintenance
YelpArchive/dataset-examples
A collection of Python example scripts from Yelp demonstrating how to work with the Yelp Academic/Open Dataset, including a JSON-to-CSV con…
101260maintenance
kakaobrain/coyo-dataset
COYO-700M is a large-scale dataset of 747M image-text pairs scraped from CommonCrawl HTML documents, with extensive image- and text-level f…
101255maintenance
onebirdrocks/geektime-ELK
Companion repository for a Chinese Geektime video course on Elasticsearch core technology and the ELK stack, containing course outlines, ex…
321210maintenance
datasets/covid-19
A cleaned and normalized time series dataset of COVID-19 confirmed cases, deaths, and recoveries, disaggregated by country and subregion, s…
651164maintenance
Lyken17/Efficient-PyTorch
A collection of best practices and example code for efficiently training large datasets like ImageNet with PyTorch. It demonstrates techniq…
321103maintenance
udacity/DSND_Term2
Udacity's Data Scientist Nanodegree Term 2 repository containing tutorial Jupyter notebooks and project templates. It covers the data scien…
321095maintenance
alanchn31/Data-Engineering-Projects
A collection of personal data engineering projects completed as part of Udacity's Data Engineering Nanodegree, covering ETL into Postgres, …
321030maintenance
KeithGalli/Pandas-Data-Science-Tasks
A collection of Jupyter Notebooks solving real-world data science tasks with Python's Pandas and Matplotlib libraries, accompanying a YouTu…
321011maintenance
robertmartin8/MachineLearningStocks
A Python tutorial project and template applying scikit-learn classifiers to historical stock prices and fundamentals to predict which stock…
321961abandoned
Made With ML
Made With ML is an open-source course teaching how to design, develop, deploy, and iterate on production-grade machine learning application…
6649233active
stefan-jansen/machine-learning-for-trading
Companion code repository for the book 'Machine Learning for Trading, 3rd Edition' by Stefan Jansen, containing 446+ Jupyter notebooks acro…
8620669active
deviantony/docker-elk
A Docker Compose template that spins up the full Elastic stack (Elasticsearch, Logstash, Kibana) with minimal configuration. It is designed…
8018388active
krishnaik06/The-Grand-Complete-Data-Science-Materials
A curated collection of video playlists, tutorials, and learning materials covering the full data science journey, from Python and statisti…
289064active
jamwithai/production-agentic-rag-course
A learner-focused course repository that walks through building a production-grade RAG system (an arXiv paper curator) week by week, coveri…
658499active
MrMimic/data-scientist-roadmap
A collection of Jupyter Notebook tutorials organized around the famous 'Road to Data Scientist' metro-map curriculum. It guides learners pr…
357386active
wassupjay/n8n-free-templates
A curated collection of 200+ ready-to-import n8n workflow JSON templates combining classic automation with AI components like vector databa…
346165active
hugo2046/QuantsPlaybook
A collection of Jupyter Notebook reproductions of 100+ quantitative investment strategies from Chinese brokerage financial engineering rese…
695892active
PacktPublishing/LLM-Engineers-Handbook
The official companion repository for the book 'LLM Engineer's Handbook' by Paul Iusztin and Maxime Labonne, containing Python code for bui…
605296active
esbatmop/MNBVC
MNBVC is a massive, continuously growing open-source Chinese text corpus aiming to rival the scale of data used to train ChatGPT, including…
754267active
oxnr/awesome-bigdata
A curated awesome-list of big data frameworks, databases, tools, papers, books, and resources. It serves as a reference catalog covering RD…
7514588maintenance
decodingai-magazine/second-brain-ai-assistant-course
An open-source course by Decoding AI that teaches building a production-ready 'Second Brain' AI assistant using agentic RAG, LLMs, fine-tun…
553058active
krishnaik06/Complete-Data-Science-With-Machine-Learning-And-NLP-2024
A companion repository for Krish Naik's Udemy course on complete data science, machine learning, NLP, and MLOps with end-to-end projects. I…
242856active
hi-weijun/PythonDataScience-Collections
A curated Chinese-language collection of links and resources for Python data analysis, covering Python basics, web scraping, visualization,…
692795active
wzhe06/SparrowRecSys
SparrowRecSys is an open-source movie recommendation system implemented as a mixed Java/Scala/Python project combining TensorFlow, Spark, a…
322776stable
elastic/examples
A collection of hands-on examples and Jupyter Notebook tutorials for the Elastic Stack (Elasticsearch, Kibana, Logstash), covering common d…
102649active
OpenCoder-llm/OpenCoder-llm
OpenCoder is a fully open and reproducible family of code large language models (1.5B and 8B base and chat variants) trained on 2.5 trillio…
222111active
big-data-europe/docker-spark
A set of Dockerfiles and Docker Compose configurations for building Apache Spark Docker images that run a standalone Spark cluster with one…
672050active
krishnaik06/Data-Science-Projects-For-Resumes
A curated collection of end-to-end data science projects (machine learning, deep learning, NLP, computer vision) designed to build a resume…
251803active

page 1 / 2 next →