# jadianes/spark-py-notebooks

Apache Spark & Python (pySpark) tutorials for Big Data Analysis and Machine Learning as IPython / Jupyter notebooks

Repository: https://github.com/jadianes/spark-py-notebooks
Canonical: https://ross.abutalabs.com/products/spark-py-notebooks
Homepage: http://jadianes.github.io/spark-py-notebooks
Language: Jupyter Notebook
License: NOASSERTION
License Family: other
Topics: spark, python, pyspark, data-analysis, mllib, ipython-notebook, notebook, ipython, data-science, machine-learning, big-data, bigdata
Last push: 2024-03-16T21:39:19+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 4137, "days_push": 900, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1659, forks 904 (observed 2026-08-28T04:05:18.305003+00:00)

## What it is
A collection of Jupyter/IPython notebooks teaching Apache Spark with Python (pySpark), from RDD basics to MLlib machine learning. It uses the KDD Cup 1999 dataset and is meant to be run in pySpark mode with a local Spark installation.

## Use cases
- learn pyspark from scratch
- understand RDD transformations like map and filter
- tutorials for machine learning with Spark MLlib
- big data analysis examples in Python
- hands-on Spark notebooks to follow along
- introduction to Spark for data science

## When to choose
- you want a free, notebook-based introduction to Spark with Python
- you learn best by running code step by step
- you need beginner-to-intermediate coverage of RDDs and MLlib

## When to avoid
- you need up-to-date Spark APIs like DataFrame-first Spark 3.x patterns
- you want production-ready code rather than tutorials
- you prefer Scala or R for Spark

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: data-science, machine-learning, etl
- domain: big-data, data-science, machine-learning, tutorials
- platform: python, cross-platform
- tags: spark, pyspark, jupyter-notebooks, mllib, rdd, big-data, educational

## Member repositories
- jadianes/spark-py-notebooks (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:18.305003+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:44:42.369011+00:00, confidence not recorded.
  - readme: https://github.com/jadianes/spark-py-notebooks (fetched 2026-08-28T04:05:18.305003+00:00, sha 67349ac210e8)
  - homepage: http://jadianes.github.io/spark-py-notebooks (fetched 2026-08-29T11:17:10.905331+00:00, sha a4e53c95b2d6)
- Data as of 2026-08-30T08:39:29.467469+00:00.
