# mahmoudparsian/data-algorithms-book

MapReduce, Spark, Java, and Scala for Data Algorithms Book

Repository: https://github.com/mahmoudparsian/data-algorithms-book
Canonical: https://ross.abutalabs.com/products/data-algorithms-book
Homepage: http://mapreduce4hackers.com
Language: Java
License: NOASSERTION
License Family: other
Topics: hadoop-mapreduce, java, distributed-computing, scala, mapreduce, data-algorithms, python, machine-learning, pyspark, distributed-algorithms, mappers, reducers, apache-hadoop, apache-spark, design-patterns, partitioning
Last push: 2024-10-14T00:19:15+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 4410, "days_push": 689, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1082, forks 652 (observed 2026-08-28T04:03:31.076064+00:00)

## What it is
The official source code repository for Mahmoud Parsian's O'Reilly books 'Data Algorithms' and related titles, containing MapReduce, Spark, Java, Scala, and PySpark examples for distributed algorithms. It serves as a companion codebase with runnable examples, build scripts, and bonus chapters.

## Use cases
- learn mapreduce with hadoop and spark examples
- find pyspark distributed algorithm code samples
- study spark machine learning algorithm implementations
- get source code for the data algorithms book
- learn mapreduce design patterns in java and scala
- examples of submitting spark jobs from java

## When to choose
- you are reading the Data Algorithms or PySpark Algorithms books and want the accompanying code
- you want worked examples of MapReduce and Spark distributed algorithms in Java, Scala, or Python
- you are learning big data processing with Hadoop and Spark through hands-on examples

## When to avoid
- you need a production-ready distributed computing library rather than educational example code
- you want a maintained framework with API stability guarantees
- you need solutions for a specific Spark version other than the one the examples target

## Facets
- artifact type: learning-resource
- maturity: active
- function: machine-learning, data-science, etl, developer-tools
- domain: big-data, data-science, microservices, tutorials
- platform: jvm, python, cross-platform
- tags: mapreduce, apache-spark, pyspark, hadoop, scala, distributed-algorithms, book-source-code, design-patterns

## Member repositories
- mahmoudparsian/data-algorithms-book (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:31.076064+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:51:16.628586+00:00, confidence not recorded.
  - readme: https://github.com/mahmoudparsian/data-algorithms-book (fetched 2026-08-28T04:03:31.076064+00:00, sha e69496d3edc6)
  - homepage: http://mapreduce4hackers.com (fetched 2026-08-29T12:53:25.124618+00:00, sha 2a982d55672e)
- Data as of 2026-08-30T08:39:29.467469+00:00.
