# NanoNets/docext

An on-premises, OCR-free unstructured data extraction, markdown conversion and benchmarking toolkit. (https://idp-leaderboard.org/)

Repository: https://github.com/NanoNets/docext
Canonical: https://ross.abutalabs.com/products/docext
Homepage: https://nanonets.com/document-parsing-and-extraction
Language: Python
License: Apache-2.0
License Family: permissive
Topics: document, document-analysis, extraction, llms, machine-learning, nlp, ocr, rag, unstructured-data, vlms, onprem, document-data-extraction, ocr-onpremise, llm-ocr, onprem-ocr, onprem-vision, onpremise, table-extraction, document-information-extraction, ocr-benchmark
Last push: 2026-03-17T09:41:55+00:00

## Health v2 (maintenance only)
Score: 50/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 72, release rhythm 28, longevity 37
- inputs: {"age_days": 526, "days_push": 169, "days_rel": 429, "gap_med": 42.5, "n_releases_24m": 3}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2085, forks 155 (observed 2026-08-28T04:06:11.908367+00:00)

## What it is
docext is an on-premises document intelligence toolkit powered by vision-language models, offering OCR-free structured data extraction, PDF/image-to-markdown conversion, and a benchmarking leaderboard for document processing tasks. It is a Python library (pip-installable) from NanoNets that also ships the Nanonets-OCR-s model for image-to-markdown conversion.

## Use cases
- extract structured fields from invoices and passports without OCR
- convert PDFs and scanned images to markdown with tables and LaTeX equations
- detect signatures and watermarks in documents
- benchmark vision-language models on document extraction tasks
- run document data extraction fully on-premises for privacy
- extract tables from unstructured documents with confidence scores

## When to choose
- you need on-premises/private document extraction without sending data to cloud APIs
- you want OCR-free structured extraction using vision-language models
- you need PDF/image to markdown conversion with semantic tagging
- you want to benchmark VLMs on IDP tasks like KIE and table extraction

## When to avoid
- you need a lightweight traditional OCR engine like Tesseract
- you lack GPU resources for running vision-language models
- you only need simple text extraction from digital PDFs
- you need a managed cloud service with SLAs

## Facets
- artifact type: library
- maturity: active
- function: ocr, nlp, machine-learning, pdf, benchmarking, rag
- domain: machine-learning, pdf, developer-tools, artificial-intelligence
- platform: python, self-hosted, cross-platform
- tags: vision-language-models, document-intelligence, information-extraction, markdown-conversion, on-premises, key-information-extraction, table-extraction, idp-leaderboard, natural-language-processing

## Member repositories
- NanoNets/docext (main) score 50

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:11.908367+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:55:49.783846+00:00, confidence not recorded.
  - readme: https://github.com/NanoNets/docext (fetched 2026-08-28T04:06:11.908367+00:00, sha 91b56618a340)
  - homepage: https://nanonets.com/document-parsing-and-extraction (fetched 2026-08-29T10:35:46.539658+00:00, sha 4a900fd45d48)
  - registry_pypi: https://pypi.org/pypi/docext/json (fetched 2026-08-29T10:35:46.548441+00:00, sha 3068d093752b)
- Data as of 2026-08-30T08:39:29.467469+00:00.
