# alanchn31/Data-Engineering-Projects

Personal Data Engineering Projects

Repository: https://github.com/alanchn31/Data-Engineering-Projects
Canonical: https://ross.abutalabs.com/products/data-engineering-projects
Language: Jupyter Notebook
License Family: other
Topics: data-lake, ingest-data, data-engineering, data-warehouse, cassandra, aws-redshift, mongodb, scrapy, spark, airflow, postgres, data-engineering-nanodegree, star-schema, data-modeling
Last push: 2023-02-08T00:44:31+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2326, "days_push": 1303, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1030, forks 208 (observed 2026-08-28T04:03:17.909097+00:00)

## What it is
A collection of personal data engineering projects completed as part of Udacity's Data Engineering Nanodegree, covering ETL into Postgres, Cassandra, MongoDB, Redshift, and S3 data lakes with Spark and Airflow. It includes course notes and example notebooks demonstrating data modeling, warehousing, and pipeline orchestration.

## Use cases
- learn data engineering through hands-on projects
- see examples of ETL pipelines with Airflow
- understand star schema data modeling in Postgres
- learn how to build a data lake with Spark and S3
- practice web scraping with Scrapy into MongoDB
- study for a data engineering nanodegree
- see how to load data into AWS Redshift

## When to choose
- you want worked examples of common data engineering patterns
- you are following or reviewing Udacity's data engineering curriculum
- you need reference code for ETL with Spark, Airflow, Cassandra, or Redshift

## When to avoid
- you need production-ready, maintained software with a license
- you want a reusable library rather than educational notebooks
- you need up-to-date tooling, as the projects were last updated in early 2023

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: etl, database, web-scraping, data-science, workflow-automation
- domain: databases, big-data, tutorials
- platform: python, cloud
- tags: data-warehouse, data-lake, airflow, spark, cassandra, redshift, mongodb, star-schema, udacity-nanodegree, jupyter-notebook, data-engineering

## Member repositories
- alanchn31/Data-Engineering-Projects (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:17.909097+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:07:32.180879+00:00, confidence not recorded.
  - readme: https://github.com/alanchn31/Data-Engineering-Projects (fetched 2026-08-28T04:03:17.909097+00:00, sha cc89ea6ede76)
- Data as of 2026-08-30T08:39:29.467469+00:00.
