# The-Japan-DataScientist-Society/100knocks-preprocess

データサイエンス100本ノック（構造化データ加工編）

Repository: https://github.com/The-Japan-DataScientist-Society/100knocks-preprocess
Canonical: https://ross.abutalabs.com/products/100knocks-preprocess
Language: HTML
License: NOASSERTION
License Family: other
Last push: 2026-05-19T21:43:51+00:00

## Health v2 (maintenance only)
Score: 60/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 83, release rhythm 8, longevity 100
- inputs: {"age_days": 2288, "days_push": 106, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2532, forks 397 (observed 2026-08-28T04:06:58.971178+00:00)

## What it is
A collection of 100 structured data processing exercises (Data Science 100 Knocks) from the Japan Data Scientist Society, with practice problems and answers in SQL, Python, and R. It ships with Docker-based environment setup, dummy supermarket purchase and personal data, and Jupyter notebooks runnable locally or on SageMaker Studio Lab / Colab.

## Use cases
- practice data preprocessing exercises in SQL, Python, and R
- learn data wrangling with hands-on problems
- set up a data science practice environment with Docker
- train data scientists in a company or university course
- practice pandas and SQL data manipulation on dummy retail data

## When to choose
- you want structured, graded practice problems for data preprocessing
- you need a ready-made Docker environment with sample data for SQL/Python/R practice
- you are teaching or self-studying data science fundamentals

## When to avoid
- you need a production data processing tool rather than exercises
- you want advanced machine learning or deep learning content
- you cannot use Docker or cloud notebooks and need a zero-setup option

## Facets
- artifact type: learning-resource
- maturity: active
- function: data-science, etl, testing
- domain: data-science, education, tutorials
- platform: python, cross-platform, cloud
- tags: sql-exercises, r, jupyter-notebooks, data-wrangling, japanese, practice-problems, docker

## Member repositories
- The-Japan-DataScientist-Society/100knocks-preprocess (main) score 60

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:58.971178+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:25:17.592980+00:00, confidence not recorded.
  - readme: https://github.com/The-Japan-DataScientist-Society/100knocks-preprocess (fetched 2026-08-28T04:06:58.971178+00:00, sha aef627c0e7e2)
- Data as of 2026-08-30T08:39:29.467469+00:00.
