# philipperemy/tensorflow-1.4-billion-password-analysis

Deep Learning model to analyze a large corpus of clear text passwords.

Repository: https://github.com/philipperemy/tensorflow-1.4-billion-password-analysis
Canonical: https://ross.abutalabs.com/products/tensorflow-14-billion-password-analysis
Language: Python
License Family: other
Topics: tensorflow, deep-learning, natural-language-processing
Last push: 2021-06-29T21:11:37+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 3184, "days_push": 1891, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1952, forks 389 (observed 2026-08-28T04:05:58.384834+00:00)

## What it is
A research project and tooling for analyzing the 1.4 billion clear-text password BreachCompilation dataset using deep learning and NLP with TensorFlow. It includes scripts to process the corpus and train generative models to study password patterns and how users modify passwords over time.

## Use cases
- analyze large leaked password corpora
- train a generative model of passwords
- study how people mutate passwords over time
- compute edit-distance pairs between passwords
- security research on password strength patterns

## When to choose
- you have the BreachCompilation dataset and want to run NLP/deep-learning analysis on it
- you are researching password generation or mutation patterns for security studies

## When to avoid
- you need a production password auditing tool for your organization
- you want a maintained library with a license - the repo has no license and last released in 2021
- you cannot legally obtain the leaked credential dataset

## Facets
- artifact type: dataset
- maturity: maintenance
- function: machine-learning, nlp, deep-learning, data-science
- domain: security, machine-learning, privacy
- platform: python
- tags: password-analysis, tensorflow, breach-compilation, generative-model, security-research, natural-language-processing, linux, macos

## Member repositories
- philipperemy/tensorflow-1.4-billion-password-analysis (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:58.384834+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:06:47.635744+00:00, confidence not recorded.
  - readme: https://github.com/philipperemy/tensorflow-1.4-billion-password-analysis (fetched 2026-08-28T04:05:58.384834+00:00, sha b68c28f41a01)
- Data as of 2026-08-30T08:39:29.467469+00:00.
