# tstanislawek/awesome-document-understanding

A curated list of resources for Document Understanding (DU) topic

Repository: https://github.com/tstanislawek/awesome-document-understanding
Canonical: https://ross.abutalabs.com/products/awesome-document-understanding
License Family: other
Topics: awesome-list, machine-learning, information-extraction, key-information-extraction, document-understanding, robotic-process-automation, document-analysis, document-layout-analysis, ocr, natural-language-processing, deep-learning, nlp, awesome, pdf, rpa, pdf-documents, document-intelligence, unstructured-data, intelligent-processing, document-ai
Last push: 2023-06-02T03:56:06+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 1975, "days_push": 1188, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1535, forks 178 (observed 2026-08-28T04:04:59.791152+00:00)

## What it is
A curated awesome-list of research papers, datasets, tools, and resources for Document Understanding (DU) and Intelligent Document Processing (IDP). It covers topics like key information extraction, document layout analysis, document question answering, and OCR for visually rich documents.

## Use cases
- find datasets for key information extraction from invoices
- learn about document layout analysis models
- discover OCR research papers and tools
- find resources for document question answering
- research intelligent document processing for RPA
- find PDF processing tools for document AI
- explore pre-training datasets for document understanding models

## When to choose
- you are researching document AI and need a starting point of papers and datasets
- you want to survey the state of the art in key information extraction or OCR
- you need to find datasets for training document understanding models

## When to avoid
- you need a working library or tool rather than a list of resources
- you need production-ready document processing software
- you need actively maintained code, since this is a curated link collection

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: ocr, nlp, machine-learning, pdf, data-science
- domain: machine-learning, pdf, awesome-lists, artificial-intelligence
- platform: -
- tags: awesome-list, document-understanding, key-information-extraction, document-layout-analysis, document-ai, intelligent-document-processing, curated-resources, natural-language-processing

## Member repositories
- tstanislawek/awesome-document-understanding (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:59.791152+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:31:10.524259+00:00, confidence not recorded.
  - readme: https://github.com/tstanislawek/awesome-document-understanding (fetched 2026-08-28T04:04:59.791152+00:00, sha 2fc09e4939d0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
