NanoNets/docext
An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/) observed · 2026-08-28
Health v2 · maintenance only
50/100
- Activity 72
- Release rhythm 28
- Longevity 37
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 42.5
- age_days: 526
- days_rel: 429
- days_push: 169
- n_releases_24m: 3
Adoption not part of the score
2085 stars · 155 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
docext is an on-premises document intelligence toolkit powered by vision-language models, offering OCR-free structured data extraction, PDF/image-to-markdown conversion, and a benchmarking leaderboard for document processing tasks. It is a Python library (pip-installable) from NanoNets that also ships the Nanonets-OCR-s model for image-to-markdown conversion.
Use cases
- extract structured fields from invoices and passports without OCR
- convert PDFs and scanned images to markdown with tables and LaTeX equations
- detect signatures and watermarks in documents
- benchmark vision-language models on document extraction tasks
- run document data extraction fully on-premises for privacy
- extract tables from unstructured documents with confidence scores
When to choose
- you need on-premises/private document extraction without sending data to cloud APIs
- you want OCR-free structured extraction using vision-language models
- you need PDF/image to markdown conversion with semantic tagging
- you want to benchmark VLMs on IDP tasks like KIE and table extraction
When to avoid
- you need a lightweight traditional OCR engine like Tesseract
- you lack GPU resources for running vision-language models
- you only need simple text extraction from digital PDFs
- you need a managed cloud service with SLAs
Facets
library · maturity active
ocr nlp machine-learning pdf benchmarking rag machine-learning pdf developer-tools artificial-intelligence python self-hosted cross-platform vision-language-models document-intelligence information-extraction markdown-conversion on-premises key-information-extraction table-extraction idp-leaderboard natural-language-processing
3 sources
- readme: https://github.com/NanoNets/docext · fetched 2026-08-28 · 91b56618a340
- homepage: https://nanonets.com/document-parsing-and-extraction · fetched 2026-08-29 · 4a900fd45d48
- registry_pypi: https://pypi.org/pypi/docext/json · fetched 2026-08-29 · 3068d093752b
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| NanoNets/docext | main | 50 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem