# datalab-to/chandra

OCR model that handles complex tables, forms, handwriting with full layout.

Repository: https://github.com/datalab-to/chandra
Canonical: https://ross.abutalabs.com/products/chandra
Homepage: https://www.datalab.to
Language: Python
License: Apache-2.0
License Family: permissive
Topics: ai, ocr
Last push: 2026-06-26T10:26:47+00:00

## Health v2 (maintenance only)
Score: 71/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 89, release rhythm 75, longevity 23
- inputs: {"age_days": 329, "days_push": 68, "days_rel": 168, "gap_med": 0, "n_releases_24m": 8}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 12171, forks 1236 (observed 2026-08-28T04:10:51.908902+00:00)

## What it is
Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSON while preserving full layout information. It supports 90+ languages and excels at complex tables, forms, math, and handwriting, with local (HuggingFace) and remote (vLLM server) inference modes.

## Use cases
- convert scanned pdfs to markdown
- extract tables from document images
- transcribe handwritten notes
- digitize forms with checkboxes
- ocr documents in multiple languages
- convert documents to structured  with layout
- extract math equations from images

## When to choose
- you need layout-aware OCR output in HTML/Markdown/JSON
- documents contain complex tables, forms, math, or handwriting
- you need multilingual OCR across 90+ languages
- you want to self-host an open-weight OCR model with vLLM

## When to avoid
- you need only plain text extraction from simple documents
- commercial self-hosting without licensing arrangements
- you lack GPU resources for local inference and don't want a remote API

## Facets
- artifact type: library
- maturity: active
- function: ocr, machine-learning, pdf, nlp
- domain: artificial-intelligence, computer-vision, pdf
- platform: python, cli
- tags: document-intelligence, layout-preservation, handwriting-recognition, table-extraction, multilingual, vllm, markdown-conversion, natural-language-processing, gpu, docker

## Member repositories
- datalab-to/chandra (main) score 71

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:51.908902+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:14:52.755030+00:00, confidence not recorded.
  - readme: https://github.com/datalab-to/chandra (fetched 2026-08-28T04:10:51.908902+00:00, sha db05f20a2067)
  - homepage: https://www.datalab.to (fetched 2026-08-29T08:12:00.014835+00:00, sha 2bbd3c9c83e2)
- Data as of 2026-08-30T08:39:29.467469+00:00.
