Ross ROSS = Recommend OSS · open-source software intelligence for agents

ttttccxxui/DataInfra-RedactionEverything

DataInfra Series. Redact EVERYTHING with local llms and vlms. observed · 2026-09-03

github.com/ttttccxxui/DataInfra-RedactionEverything · Python · NOASSERTION (other) observed · 2026-09-03

Health v2 · maintenance only

60/100

  • Activity 100
  • Release rhythm 35
  • Longevity 15

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 217
  • days_rel: n/a
  • days_push: 0
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1147 stars · 169 forks observed · 2026-09-03

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and plain text using local LLMs, VLMs, OCR, and semantic NER. It provides human review, batch processing, configurable industry schemas, and export workflows without sending raw documents to remote APIs.

Use cases

  • redact pii from pdfs locally
  • anonymize scanned documents with ocr
  • remove sensitive data from word files and images
  • batch redact documents without cloud apis
  • detect names ids and signatures in documents
  • review and export redacted files
  • self-hosted document anonymization

When to choose

  • you need local-only processing so raw documents never leave your machine
  • you need to redact both text and visual elements like seals, faces, and signatures
  • you want human-in-the-loop review before finalizing redactions
  • you process scanned or image-based documents, not just plain text

When to avoid

  • you need a tool for commercial or production use without negotiating a separate license
  • you want a lightweight rule-based PII scanner rather than a full workbench with GPU model dependencies
  • you cannot run local LLM/VLM models due to limited GPU memory
  • you need permissive open-source licensing for redistribution or hosted services

Facets

application · maturity active

ocr nlp image-processing pdf privacy security machine-learning llm-inference privacy computer-vision pdf files self-hosted artificial-intelligence python self-hosted data-redaction pii-detection document-anonymization local-llm vlm ner human-review batch-processing personal-use-license natural-language-processing docker linux gpu

1 source

Member repositories

RepositoryRoleHealth v2
ttttccxxui/DataInfra-RedactionEverythingmain60

For agents

markdown · JSON · MCP: product_card(name="ttttccxxui/DataInfra-RedactionEverything")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem