axa-group/Parsr
Transforms PDF, Documents and Images into Enriched Structured Data observed · 2026-08-28
Health v2 · maintenance only
56/100
- Activity 73
- Release rhythm 8
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 2585
- days_rel: n/a
- days_push: 166
- n_releases_24m: 0
Adoption not part of the score
6178 stars · 318 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Parsr is a document parsing and extraction toolchain that transforms PDFs, images, docx, and eml files into clean, enriched structured data in JSON, Markdown, CSV, or TXT. It offers a REST API, Python client, and optional GUI, deployable via Docker.
Use cases
- extract structured data from pdf documents
- convert scanned documents to or markdown
- parse tables and headings from pdfs
- ocr images into text data
- automate document data entry pipelines
- clean and restructure document hierarchy
When to choose
- you need structured extraction from pdfs, images, or office documents
- you want a self-hosted document parsing API with a python client
- you need table, heading, and list detection from documents
When to avoid
- you need active maintenance or security patches - the project is officially unmaintained
- you want a lightweight library rather than a dockerized API service
- you need cutting-edge LLM-based document understanding
Facets
library · maturity abandoned
ocr pdf parser nlp etl pdf files self-hosted python document-parsing pdf-extraction structured-data document-cleaning table-detection api-server unmaintained natural-language-processing data-engineering docker web-server nodejs
1 source
- readme: https://github.com/axa-group/Parsr · fetched 2026-08-28 · 5cfd4b9ff1ad
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| axa-group/Parsr | main | 56 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem