# Topdu/OpenOCR

OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

Repository: https://github.com/Topdu/OpenOCR
Canonical: https://ross.abutalabs.com/products/openocr
Language: Python
License: Apache-2.0
License Family: permissive
Topics: chineseocr, ocr, ocr-pytorch, scene-text-detection, scene-text-recognition, document-analysis, document-parsing, document-processing
Last push: 2026-08-04T04:54:39+00:00

## Health v2 (maintenance only)
Score: 58/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 96, release rhythm 8, longevity 58
- inputs: {"age_days": 824, "days_push": 29, "days_rel": 648, "gap_med": null, "n_releases_24m": 1}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1437, forks 142 (observed 2026-08-28T04:04:43.745313+00:00)

## What it is
OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and table recognition, and document parsing. It provides commercial-grade OCR systems, a unified training/evaluation benchmark, and reproductions of academic paper implementations.

## Use cases
- extract text from images with ocr
- parse documents into structured text
- recognize tables in scanned pdfs
- detect and recognize scene text in photos
- convert formulas in images to latex
- benchmark ocr models on standard datasets
- train custom text recognition models in pytorch
- chinese and english ocr

## When to choose
- you need a full OCR pipeline including detection, recognition, and document parsing
- you want reproducible implementations of OCR research papers for benchmarking
- you need commercial-grade OCR with lightweight deployable models
- you work with Chinese and English text recognition

## When to avoid
- you only need speech-to-text rather than visual text recognition
- you need handwriting-heavy or low-resource language OCR beyond the toolkit's supported models
- you want a fully managed cloud OCR API without running models yourself

## Facets
- artifact type: library
- maturity: active
- function: ocr, machine-learning, deep-learning, image-processing, pdf
- domain: computer-vision, artificial-intelligence, pdf
- platform: python, windows
- tags: text-detection, text-recognition, document-parsing, table-recognition, formula-recognition, scene-text, pytorch, research-benchmark, natural-language-processing, linux, macos, gpu

## Member repositories
- Topdu/OpenOCR (main) score 58

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:43.745313+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:36:42.794759+00:00, confidence not recorded.
  - readme: https://github.com/Topdu/OpenOCR (fetched 2026-08-28T04:04:43.745313+00:00, sha 2d7a60f71be5)
- Data as of 2026-08-30T08:39:29.467469+00:00.
