Topdu/OpenOCR
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers. observed · 2026-08-28
Health v2 · maintenance only
58/100
- Activity 96
- Release rhythm 8
- Longevity 58
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 824
- days_rel: 648
- days_push: 29
- n_releases_24m: 1
Adoption not part of the score
1437 stars · 142 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and table recognition, and document parsing. It provides commercial-grade OCR systems, a unified training/evaluation benchmark, and reproductions of academic paper implementations.
Use cases
- extract text from images with ocr
- parse documents into structured text
- recognize tables in scanned pdfs
- detect and recognize scene text in photos
- convert formulas in images to latex
- benchmark ocr models on standard datasets
- train custom text recognition models in pytorch
- chinese and english ocr
When to choose
- you need a full OCR pipeline including detection, recognition, and document parsing
- you want reproducible implementations of OCR research papers for benchmarking
- you need commercial-grade OCR with lightweight deployable models
- you work with Chinese and English text recognition
When to avoid
- you only need speech-to-text rather than visual text recognition
- you need handwriting-heavy or low-resource language OCR beyond the toolkit's supported models
- you want a fully managed cloud OCR API without running models yourself
Facets
library · maturity active
ocr machine-learning deep-learning image-processing pdf computer-vision artificial-intelligence pdf python windows text-detection text-recognition document-parsing table-recognition formula-recognition scene-text pytorch research-benchmark natural-language-processing linux macos gpu
1 source
- readme: https://github.com/Topdu/OpenOCR · fetched 2026-08-28 · 2d7a60f71be5
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| Topdu/OpenOCR | main | 58 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem