resource: etl
121 resources, primary matches first, then adoption-weighted; health v2 shown.
| Resource | Health v2 | Stars | Maturity |
|---|---|---|---|
| DataTalksClub/data-engineering-zoomcamp A free 9-week open-source course teaching data engineering fundamentals by building an end-to-end production-ready data pipeline. It includ… | 77 | 45005 | active |
| andkret/Cookbook The Data Engineering Cookbook is a free, open-source book by Andreas Kretz covering the skills, tools, and best practices needed to become … | 75 | 15371 | active |
| zhisheng17/flink-learning A comprehensive Apache Flink learning repository with example code covering DataStream API, Table API & SQL, connectors, metrics, and produ… | 69 | 15095 | active |
| DataTalksClub/machine-learning-zoomcamp A free 4-month cohort-based course from DataTalks.Club teaching machine learning engineering, covering the full path from problem framing a… | 68 | 14031 | active |
| joevess/IPTV An automatically updated collection of IPTV live-stream playlists (m3u8) aggregating sources from haoqu, TVBox, and other public sources, s… | 28 | 10276 | active |
| igorbarinov/awesome-data-engineering A curated awesome-list of data engineering tools, databases, frameworks, and resources for software developers. It organizes links across c… | 74 | 8992 | active |
| rfordatascience/tidytuesday TidyTuesday is a weekly social data project that publishes real-world datasets for people to practice data wrangling, visualization, and mo… | 67 | 8359 | active |
| pditommaso/awesome-pipeline A curated awesome-list cataloging pipeline toolkits, workflow engines, and data orchestration frameworks. It serves as a discovery resource… | 75 | 6624 | active |
| togethercomputer/RedPajama-Data RedPajama-Data provides code and pipelines for building RedPajama-V2, an open dataset with over 30 trillion tokens of web text for training… | 68 | 4980 | active |
| mlabonne/llm-datasets A curated list of datasets and tools for post-training large language models, covering supervised fine-tuning, preference, and reasoning da… | 60 | 4757 | active |
| tensorflow/datasets TensorFlow Datasets (TFDS) is a curated collection of ready-to-use public datasets exposed as tf.data.Dataset objects for TensorFlow, JAX, … | 85 | 4581 | stable |
| decodingai-magazine/llm-twin-course A free hands-on course (source code plus 12 lessons) that teaches how to build an end-to-end production-ready LLM and RAG system by creatin… | 60 | 4385 | active |
| electricitymaps/electricitymaps-contrib A community-maintained collection of Python parsers that collect and standardize electricity data (production, exchanges, prices) from offi… | 89 | 4025 | active |
| jghoman/awesome-apache-airflow A curated awesome-list of resources about Apache Airflow, the workflow orchestration platform. It collects tutorials, deployment solutions,… | 69 | 3925 | active |
| Jon-Becker/prediction-market-analysis A Python framework for collecting and analyzing prediction market data from Polymarket and Kalshi, including the largest publicly available… | 60 | 3774 | active |
| xyflow/awesome-node-based-uis A curated awesome-list of resources, libraries, and tools for building node-based UIs (node editors, visual programming environments, and w… | 45 | 3664 | active |
| pawl/awesome-etl A curated awesome-list of ETL (extract, transform, load) frameworks, libraries, and software, favoring mainstream, well-supported open-sour… | 68 | 3585 | active |
| MIT-LCP/mimic-code A community-maintained repository of code for building, loading, and analyzing the MIMIC family of critical care databases (MIMIC-III, MIMI… | 79 | 3355 | active |
| openaddresses/openaddresses OpenAddresses is a global, openly licensed collection of address, cadastral parcel, building footprint, and street centerline data sources.… | 77 | 3242 | active |
| igrigorik/gharchive.org GH Archive records the public GitHub timeline into hourly JSON event archives and makes them freely downloadable, plus available as a publi… | 43 | 3072 | active |
| GoogleCloudPlatform/professional-services A collection of example solutions, tools, and reference architectures developed by Google Cloud's Professional Services team, primarily in … | 76 | 3065 | active |
| danielbeach/data-engineering-practice A collection of hands-on data engineering practice problems covering Python data processing, file formats, SQL, Postgres, PySpark, and data… | 32 | 2845 | active |
| The-Japan-DataScientist-Society/100knocks-preprocess A collection of 100 structured data processing exercises (Data Science 100 Knocks) from the Japan Data Scientist Society, with practice pro… | 60 | 2532 | active |
| binance/binance-public-data A repository documenting and providing access to Binance's public historical market data, downloadable as daily or monthly files for spot, … | 32 | 2462 | active |
| thenaturalist/awesome-business-intelligence A curated awesome-list of business intelligence tools, split into SaaS and open-source categories covering analytics clients, data visualiz… | 32 | 2314 | active |
| toddwschneider/nyc-taxi-data A collection of R and shell scripts to download, process, and import 3+ billion NYC taxi and for-hire vehicle (Uber, Lyft) trip records int… | 49 | 2081 | active |
| confluentinc/examples A curated collection of demos and examples showcasing Apache Kafka, Apache Flink, and Confluent Platform event streaming, including Conflue… | 77 | 2064 | active |
| awslabs/open-data-registry A registry of publicly available datasets that can be accessed via AWS resources, with each dataset described in a YAML metadata file. It p… | 77 | 1774 | active |
| scrollmapper/bible_databases A collection of 140 Bible translations distributed as databases in multiple formats including MySQL, SQLite, CSV, JSON, YAML, TXT, Markdown… | 73 | 1675 | active |
| oxylabs/scrape-google-python A Python tutorial repository from Oxylabs demonstrating how to scrape Google search results (SERPs) using Oxylabs' SERP Scraper API. It inc… | 63 | 1629 | active |
| AlexTheAnalyst/PortfolioProjects A collection of data analyst portfolio projects with code and queries from Alex The Analyst's tutorials, primarily in Jupyter Notebooks. It… | 32 | 1568 | active |
| aws-samples/aws-glue-samples A collection of code samples, tutorials, and utilities demonstrating AWS Glue, Amazon's serverless data integration and ETL service. It cov… | 76 | 1537 | active |
| Marvomatic/n8n-templates A collection of free and premium n8n workflow templates for automating SEO, content optimization, and data analysis tasks. Templates integr… | 44 | 1537 | active |
| datawhalechina/hands-on-data-analysis An open-source Chinese-language hands-on course from Datawhale that teaches data analysis through project-based Jupyter notebooks covering … | 32 | 1524 | active |
| duneanalytics/spellbook Spellbook is a community-maintained collection of dbt SQL models that define curated, decoded datasets (like dex.trades and NFT data) on Du… | 67 | 1513 | active |
| TurboWay/bigdata_analyse A collection of hands-on big data analysis projects in Python, SQL, and HQL, each with documentation, datasets, and full pipelines from cle… | 32 | 5325 | maintenance |
| GoogleCloudPlatform/data-science-on-gcp Companion source code repository for the O'Reilly book 'Data Science on the Google Cloud Platform' by Valliappa Lakshmanan, provided as Jup… | 63 | 1427 | active |
| chenditc/investment_data A crowdsourced Chinese stock market dataset (A-shares) maintained in a Dolt database and exported to Microsoft Qlib binary format via Pytho… | 95 | 1423 | active |
| lvgalvao/data-engineering-roadmap The official repository of a professional training program in Data Engineering and AI (university extension) by Jornada de Dados. It contai… | 63 | 1394 | active |
| spark-examples/pyspark-examples A collection of runnable Python example scripts demonstrating PySpark RDD, DataFrame, and SQL operations, companion to the SparkByExamples … | 57 | 1365 | active |
| singer-io/getting-started The official getting-started guide and documentation for the Singer open-source ETL standard, which defines how extraction scripts (Taps) a… | 48 | 1345 | active |
| Data-Learn/data-engineering Data Learn is a free open educational resource and course series teaching data engineering and analytics, covering BI tools, databases, ETL… | 40 | 1332 | active |
| datascale-ai/data_engineering_book An open-source book (with GitHub Pages site and runnable Python code) on data engineering for large language models, covering pretraining d… | 65 | 1300 | active |
| mahmoudparsian/pyspark-tutorial A tutorial repository of Jupyter Notebook examples teaching basic distributed algorithms using PySpark, the Python API for Apache Spark. It… | 43 | 1279 | active |
| MrSuiChuan/data-warehouse-learning A Chinese-language educational project teaching real-time and offline data warehouse construction for an e-commerce system, with hands-on c… | 61 | 1218 | active |
| codemayq/chinese-chatbot-corpus A curated collection and unified processing pipeline for publicly available Chinese chit-chat conversation corpora, aggregating eight sourc… | 32 | 4194 | maintenance |
| DataExpert-io/llm-driven-data-engineering A public educational repository accompanying DataExpert's bootcamp course on LLM-driven data engineering concepts. It contains Python labs … | 28 | 1156 | active |
| databricks/learning-spark Example code accompanying the O'Reilly 'Learning Spark' book, with implementations in Java, Scala, and Python. The examples target Spark 1.… | 73 | 3893 | maintenance |
| elastic/ecs Elastic Common Schema (ECS) is an open-source specification defining a common set of fields, data types, and usage hierarchies for ingestin… | 94 | 1119 | active |
| evansiroky/timezone-boundary-builder A tool and dataset project that builds the world's timezone boundaries from OpenStreetMap data, released as shapefiles and GeoJSON. Each bo… | 89 | 1114 | active |
| mahmoudparsian/data-algorithms-book The official source code repository for Mahmoud Parsian's O'Reilly books 'Data Algorithms' and related titles, containing MapReduce, Spark,… | 32 | 1082 | active |
| apache/flink-training A collection of hands-on lab exercises and solutions accompanying the official Apache Flink training documentation. It provides a Gradle-ba… | 69 | 1048 | active |
| OpenSourceAP/CrossSection Code and data accompanying Chen and Zimmermann's paper on open source cross-sectional asset pricing. It reproduces dozens of stock-level pr… | 49 | 1034 | active |
| tomwhite/hadoop-book Example source code accompanying O'Reilly's 'Hadoop: The Definitive Guide' by Tom White, covering the Apache Hadoop ecosystem including Had… | 32 | 3498 | maintenance |
| data-science-on-aws/data-science-on-aws Companion Jupyter Notebook repository for the O'Reilly book 'Data Science on AWS', containing end-to-end AI/ML pipeline examples using Amaz… | 32 | 3433 | maintenance |
| databricks/Spark-The-Definitive-Guide The official code repository for the book 'Spark: The Definitive Guide' by Bill Chambers and Matei Zaharia, containing example code and dat… | 32 | 3144 | maintenance |
| yyx990803/build-your-own-mint A demo/tutorial project by Evan You showing how to build a personal finance analytics pipeline using Plaid bank data, Google Sheets, and Ci… | 32 | 2538 | maintenance |
| qiurunze123/threadandjuc A Java educational project teaching multithreading and concurrency from basics to advanced, culminating in a 'three-high' (high-availabilit… | 32 | 2160 | maintenance |
| AlexIoannides/pyspark-example-project A reference example project demonstrating best practices for structuring PySpark ETL jobs and applications. It shows how to organize code f… | 10 | 2119 | maintenance |
| chris1610/pbpython A collection of code, Jupyter notebooks, and examples accompanying the Practical Business Python blog (pbpython.com), which teaches Python … | 32 | 1992 | maintenance |
| DXY-COVID-19 A time-series data warehouse of COVID-19 (2019-nCoV) infection statistics for China, scraped from Dingxiangyuan (DXY) and published as CSV/… | 10 | 1971 | maintenance |
| san089/Udacity-Data-Engineering-Projects A collection of Udacity Data Engineering Nanodegree projects covering data modeling with Postgres and Cassandra, cloud data warehousing wit… | 32 | 1969 | maintenance |
| EleutherAI/the-pile The Pile is a large, diverse, open-source language modeling dataset composed of many smaller text sources combined together. This repositor… | 32 | 1673 | maintenance |
| jadianes/spark-py-notebooks A collection of Jupyter/IPython notebooks teaching Apache Spark with Python (pySpark), from RDD basics to MLlib machine learning. It uses t… | 32 | 1659 | maintenance |
| gururise/AlpacaDataCleaned A cleaned and curated version of the Stanford Alpaca instruction-tuning dataset used to train the Alpaca LLM. It fixes quality issues in th… | 62 | 1606 | maintenance |
| confluentinc/demo-scene A collection of scripts, Docker Compose setups, and sample code supporting Confluent's demos, talks, and blogs about Apache Kafka and its e… | 67 | 1565 | maintenance |
| sryza/aas Source code examples accompanying the O'Reilly book 'Advanced Analytics with Spark', written in Scala and built with Maven. It provides run… | 23 | 1525 | maintenance |
| griffithlab/rnaseq_tutorial An educational web resource and working pipeline for RNA-seq analysis on the cloud, covering file formats, reference genomes, expression an… | 23 | 1435 | maintenance |
| gtoonstra/etl-with-airflow A community-maintained collection of ETL best practices, usage patterns, and examples for Apache Airflow, published as documentation with a… | 32 | 1356 | maintenance |
| eleanorlutz/asteroids_atlas_of_space An open-source tutorial and dataset for creating a detailed map of asteroid orbits in the solar system using NASA data. It includes Python … | 32 | 1307 | maintenance |
| WillKoehrsen/machine-learning-project-walkthrough A Jupyter Notebook-based walkthrough demonstrating a complete end-to-end machine learning solution on a real-world dataset. It shows how al… | 32 | 1301 | maintenance |
| rsvp/fecon235 A collection of Jupyter notebooks for financial economics research, providing high-level Python interfaces to economic data sources like FR… | 32 | 1275 | maintenance |
| YelpArchive/dataset-examples A collection of Python example scripts from Yelp demonstrating how to work with the Yelp Academic/Open Dataset, including a JSON-to-CSV con… | 10 | 1260 | maintenance |
| kakaobrain/coyo-dataset COYO-700M is a large-scale dataset of 747M image-text pairs scraped from CommonCrawl HTML documents, with extensive image- and text-level f… | 10 | 1255 | maintenance |
| onebirdrocks/geektime-ELK Companion repository for a Chinese Geektime video course on Elasticsearch core technology and the ELK stack, containing course outlines, ex… | 32 | 1210 | maintenance |
| datasets/covid-19 A cleaned and normalized time series dataset of COVID-19 confirmed cases, deaths, and recoveries, disaggregated by country and subregion, s… | 65 | 1164 | maintenance |
| Lyken17/Efficient-PyTorch A collection of best practices and example code for efficiently training large datasets like ImageNet with PyTorch. It demonstrates techniq… | 32 | 1103 | maintenance |
| udacity/DSND_Term2 Udacity's Data Scientist Nanodegree Term 2 repository containing tutorial Jupyter notebooks and project templates. It covers the data scien… | 32 | 1095 | maintenance |
| alanchn31/Data-Engineering-Projects A collection of personal data engineering projects completed as part of Udacity's Data Engineering Nanodegree, covering ETL into Postgres, … | 32 | 1030 | maintenance |
| KeithGalli/Pandas-Data-Science-Tasks A collection of Jupyter Notebooks solving real-world data science tasks with Python's Pandas and Matplotlib libraries, accompanying a YouTu… | 32 | 1011 | maintenance |
| robertmartin8/MachineLearningStocks A Python tutorial project and template applying scikit-learn classifiers to historical stock prices and fundamentals to predict which stock… | 32 | 1961 | abandoned |
| Made With ML Made With ML is an open-source course teaching how to design, develop, deploy, and iterate on production-grade machine learning application… | 66 | 49233 | active |
| stefan-jansen/machine-learning-for-trading Companion code repository for the book 'Machine Learning for Trading, 3rd Edition' by Stefan Jansen, containing 446+ Jupyter notebooks acro… | 86 | 20669 | active |
| deviantony/docker-elk A Docker Compose template that spins up the full Elastic stack (Elasticsearch, Logstash, Kibana) with minimal configuration. It is designed… | 80 | 18388 | active |
| krishnaik06/The-Grand-Complete-Data-Science-Materials A curated collection of video playlists, tutorials, and learning materials covering the full data science journey, from Python and statisti… | 28 | 9064 | active |
| jamwithai/production-agentic-rag-course A learner-focused course repository that walks through building a production-grade RAG system (an arXiv paper curator) week by week, coveri… | 65 | 8499 | active |
| MrMimic/data-scientist-roadmap A collection of Jupyter Notebook tutorials organized around the famous 'Road to Data Scientist' metro-map curriculum. It guides learners pr… | 35 | 7386 | active |
| wassupjay/n8n-free-templates A curated collection of 200+ ready-to-import n8n workflow JSON templates combining classic automation with AI components like vector databa… | 34 | 6165 | active |
| hugo2046/QuantsPlaybook A collection of Jupyter Notebook reproductions of 100+ quantitative investment strategies from Chinese brokerage financial engineering rese… | 69 | 5892 | active |
| PacktPublishing/LLM-Engineers-Handbook The official companion repository for the book 'LLM Engineer's Handbook' by Paul Iusztin and Maxime Labonne, containing Python code for bui… | 60 | 5296 | active |
| esbatmop/MNBVC MNBVC is a massive, continuously growing open-source Chinese text corpus aiming to rival the scale of data used to train ChatGPT, including… | 75 | 4267 | active |
| oxnr/awesome-bigdata A curated awesome-list of big data frameworks, databases, tools, papers, books, and resources. It serves as a reference catalog covering RD… | 75 | 14588 | maintenance |
| decodingai-magazine/second-brain-ai-assistant-course An open-source course by Decoding AI that teaches building a production-ready 'Second Brain' AI assistant using agentic RAG, LLMs, fine-tun… | 55 | 3058 | active |
| krishnaik06/Complete-Data-Science-With-Machine-Learning-And-NLP-2024 A companion repository for Krish Naik's Udemy course on complete data science, machine learning, NLP, and MLOps with end-to-end projects. I… | 24 | 2856 | active |
| hi-weijun/PythonDataScience-Collections A curated Chinese-language collection of links and resources for Python data analysis, covering Python basics, web scraping, visualization,… | 69 | 2795 | active |
| wzhe06/SparrowRecSys SparrowRecSys is an open-source movie recommendation system implemented as a mixed Java/Scala/Python project combining TensorFlow, Spark, a… | 32 | 2776 | stable |
| elastic/examples A collection of hands-on examples and Jupyter Notebook tutorials for the Elastic Stack (Elasticsearch, Kibana, Logstash), covering common d… | 10 | 2649 | active |
| OpenCoder-llm/OpenCoder-llm OpenCoder is a fully open and reproducible family of code large language models (1.5B and 8B base and chat variants) trained on 2.5 trillio… | 22 | 2111 | active |
| big-data-europe/docker-spark A set of Dockerfiles and Docker Compose configurations for building Apache Spark Docker images that run a standalone Spark cluster with one… | 67 | 2050 | active |
| krishnaik06/Data-Science-Projects-For-Resumes A curated collection of end-to-end data science projects (machine learning, deep learning, NLP, computer vision) designed to build a resume… | 25 | 1803 | active |
page 1 / 2 next →