# ttttccxxui/DataInfra-RedactionEverything

DataInfra Series. Redact EVERYTHING with local llms and vlms.

Repository: https://github.com/ttttccxxui/DataInfra-RedactionEverything
Canonical: https://ross.abutalabs.com/products/datainfra-redactioneverything
Language: Python
License: NOASSERTION
License Family: other
Last push: 2026-09-02T06:11:28+00:00

## Health v2 (maintenance only)
Score: 60/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 100, release rhythm 35, longevity 15
- inputs: {"age_days": 217, "days_push": 0, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1147, forks 169 (observed 2026-09-03T02:15:19.995114+00:00)

## What it is
A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and plain text using local LLMs, VLMs, OCR, and semantic NER. It provides human review, batch processing, configurable industry schemas, and export workflows without sending raw documents to remote APIs.

## Use cases
- redact pii from pdfs locally
- anonymize scanned documents with ocr
- remove sensitive data from word files and images
- batch redact documents without cloud apis
- detect names ids and signatures in documents
- review and export redacted files
- self-hosted document anonymization

## When to choose
- you need local-only processing so raw documents never leave your machine
- you need to redact both text and visual elements like seals, faces, and signatures
- you want human-in-the-loop review before finalizing redactions
- you process scanned or image-based documents, not just plain text

## When to avoid
- you need a tool for commercial or production use without negotiating a separate license
- you want a lightweight rule-based PII scanner rather than a full workbench with GPU model dependencies
- you cannot run local LLM/VLM models due to limited GPU memory
- you need permissive open-source licensing for redistribution or hosted services

## Facets
- artifact type: application
- maturity: active
- function: ocr, nlp, image-processing, pdf, privacy, security, machine-learning, llm-inference
- domain: privacy, computer-vision, pdf, files, self-hosted, artificial-intelligence
- platform: python, self-hosted
- tags: data-redaction, pii-detection, document-anonymization, local-llm, vlm, ner, human-review, batch-processing, personal-use-license, natural-language-processing, docker, linux, gpu

## Member repositories
- ttttccxxui/DataInfra-RedactionEverything (main) score 60

## Provenance
- Observed fields: from GitHub, fetched 2026-09-03T02:15:19.995114+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:34:13.219799+00:00, confidence not recorded.
  - readme: https://github.com/ttttccxxui/DataInfra-RedactionEverything (fetched 2026-09-03T02:15:19.995114+00:00, sha 686ba79a9ecd)
- Data as of 2026-08-30T08:39:29.467469+00:00.
