# san089/Udacity-Data-Engineering-Projects

Few projects related to Data Engineering including Data Modeling, Infrastructure setup on cloud, Data Warehousing and Data Lake development.

Repository: https://github.com/san089/Udacity-Data-Engineering-Projects
Canonical: https://ross.abutalabs.com/products/udacity-data-engineering-projects
Language: Python
License: NOASSERTION
License Family: other
Topics: data, data-engineering, data-engineering-pipeline, etl-pipeline, cassandra-database, postgresql-database, data-modeling, data-warehouse, data-lake, airflow, airflow-operators, cluster, cassandra, infrastructure, postgres, aws, aws-ec2, aws-sdk, aws-s3, cloudformation
Last push: 2022-08-26T00:09:15+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2417, "days_push": 1469, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1969, forks 593 (observed 2026-08-28T04:06:00.759633+00:00)

## What it is
A collection of Udacity Data Engineering Nanodegree projects covering data modeling with Postgres and Cassandra, cloud data warehousing with Redshift, data lakes with Spark on EMR, and pipeline orchestration with Airflow. Each project includes Python ETL code and AWS infrastructure setup examples.

## Use cases
- learn data engineering by example
- build an ETL pipeline with Python and Postgres
- set up a data warehouse on AWS Redshift
- create a data lake with Spark on EMR
- orchestrate scheduled ETL jobs with Airflow
- practice data modeling with Cassandra
- study for a data engineering interview

## When to choose
- you want hands-on, project-based learning for data engineering concepts
- you need reference implementations of ETL pipelines on AWS
- you are following or supplementing the Udacity Data Engineering Nanodegree

## When to avoid
- you need production-ready, maintained data engineering tooling
- you want a library or framework to import into your own codebase
- you need up-to-date cloud practices, as the projects were last updated in 2022

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: etl, database, data-science, infrastructure-as-code, workflow-automation
- domain: big-data, cloud-computing, databases
- platform: python, cloud
- tags: data-modeling, data-warehouse, data-lake, apache-airflow, apache-cassandra, postgresql, amazon-redshift, spark, aws-emr, udacity, tutorial-projects, etl-pipelines, data-engineering, aws

## Member repositories
- san089/Udacity-Data-Engineering-Projects (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:00.759633+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:05:10.339205+00:00, confidence not recorded.
  - readme: https://github.com/san089/Udacity-Data-Engineering-Projects (fetched 2026-08-28T04:06:00.759633+00:00, sha 7ee1edd9a462)
- Data as of 2026-08-30T08:39:29.467469+00:00.
