function: ocr
484 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| drmingler/docling-api A self-hostable FastAPI backend service that converts documents (PDF, DOCX, PPTX, HTML, images, CSV, AsciiDoc, Markdown) into Markdown usin… | 65 | 1557 | active |
| hiroi-sora/PaddleOCR-json An offline OCR command-line executable compiled from PaddleOCR C++ that recognizes text in images and outputs results as JSON strings. It c… | 30 | 1542 | active |
| studyhelperhelper/studyhelper An Android app disguised as a Sudoku game that automates earning daily points in the Xuexi Qiangguo (学习强国) app. It uses accessibility servi… | 23 | 1533 | active |
| NanoNets/docstrange DocStrange is a Python library and tool that converts documents (PDF, DOCX, PPTX, XLSX, images, URLs) into Markdown, JSON, CSV, or HTML usi… | 40 | 1531 | active |
| WenmuZhou/PytorchOCR A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP… | 59 | 1523 | active |
| DocumindHQ/documind Documind is an open-source Node.js library that uses AI/LLMs to extract structured JSON data from PDFs and other documents based on customi… | 33 | 1521 | active |
| jonaswinkler/paperless-ng Paperless-ng is a self-hosted document management system that scans, OCRs, indexes, and archives physical documents with full-text search a… | 10 | 5414 | maintenance |
| Turing-Project/WriteGPT WriteGPT is a generative text-creation AI framework built on GPT-2 and other models (EAST, CRNN, BERT), fine-tuned to generate Chinese exam… | 23 | 5289 | maintenance |
| dadoonet/fscrawler FSCrawler is a Java-based file system crawler that indexes binary documents (PDF, MS Office, Open Office) into Elasticsearch, tracking new,… | 88 | 1450 | active |
| beyondtranslate/beyondtranslate-ce BeyondTranslate (formerly Biyi) is a cross-platform desktop translation app for macOS, Windows, and Linux built with Flutter and Rust. It c… | 67 | 1449 | active |
| serratus/quaggaJS QuaggaJS is a barcode-scanner library written entirely in JavaScript that supports real-time localization and decoding of barcode types suc… | 23 | 5207 | maintenance |
| sdcb/PaddleSharp A .NET/C# wrapper around Baidu's PaddleInference C API, providing PaddleOCR, PaddleDetection, rotation detection, Chinese segmentation, and… | 72 | 1441 | active |
| Topdu/OpenOCR OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta… | 58 | 1437 | active |
| blacklanternsecurity/MANSPIDER MANSPIDER is a Python CLI tool that crawls SMB shares across entire networks to find files by filename or content, with regex support and t… | 77 | 1406 | active |
| alibaba/Logics-Parsing Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul… | 54 | 1402 | active |
| sMythicalBird/ZenlessZoneZero-Auto A Python-based automation framework for the game Zenless Zone Zero that uses image classification, template matching, and OCR to perform au… | 22 | 1378 | active |
| zelon88/HRConvert2 HRConvert2 is a self-hosted, resource-aware file conversion server written in PHP that supports 488 file formats across documents, images, … | 99 | 1364 | active |
| saucepleez/taskt taskt (formerly sharpRPA) is a free, open-source robotic process automation (RPA) client built in C# on the .NET Framework. It provides a W… | 38 | 1363 | active |
| huridocs/pdf-document-layout-analysis A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl… | 82 | 1346 | active |
| CoderWanFeng/python-office python-office is a Python office-automation library that wraps PDF, Word, Excel, PPT, image, video, email, WeChat, and OCR operations behin… | 67 | 1344 | active |
| wormtql/yas Yas is a fast screen-scanning tool that uses a custom-trained SVTR OCR model to read Genshin Impact and Honkai: Star Rail artifact stats di… | 40 | 1336 | active |
| GauravSingh9356/J.A.R.V.I.S A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,… | 48 | 1332 | active |
| ChaokunHong/MetaScreener MetaScreener is an open-source AI tool that automates title/abstract and full-text PDF screening for systematic reviews using an ensemble o… | 71 | 1328 | active |
| deanmalmgren/textract A Python library that extracts text from virtually any document format (PDF, DOCX, PPTX, HTML, images, and more) through a simple unified i… | 93 | 4698 | maintenance |
| gavrielc/Nano-PDF A Python CLI tool that edits PDF slides using natural language prompts, powered by Google's Gemini 3 Pro Image model. It renders pages to i… | 41 | 1321 | active |
| CBIhalsen/PolyglotPDF PolyglotPDF is a Python-based multilingual eBook and PDF translation tool that preserves original layouts while translating, supporting bot… | 42 | 1315 | active |
| Intuition-Lab/personal-model Personal Model (Persome) is a local-first macOS runtime that captures focused cross-app activity and builds an evidence-linked personal mem… | 77 | 1314 | active |
| ndl-lab/ndlocr-lite NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma… | 77 | 1309 | active |
| ttop32/MouseTooltipTranslator A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports … | 95 | 1300 | active |
| wisupai/e2m E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using… | 23 | 1294 | active |
| sist2app/sist2 sist2 is a fast, multi-threaded file system indexer that extracts text, metadata, and thumbnails from common file types, with OCR support v… | 99 | 1289 | active |
| lzhgus/Capso Capso is a free, open-source native macOS app for screenshots and screen recording, built with Swift 6.0 and SwiftUI as an alternative to C… | 81 | 1286 | active |
| jbarrow/commonforms CommonForms is a Python package and CLI that uses trained object-detection models (FFDNet-S/L) to automatically detect form fields in a PDF… | 65 | 1284 | active |
| flutter-ml/google_ml_kit_flutter A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa… | 76 | 1274 | active |
| Yuliang-Liu/MonkeyOCRv2 MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2… | 58 | 1256 | active |
| unjs/unpdf unpdf is a TypeScript library providing PDF extraction and rendering utilities that work across all JavaScript runtimes, including Node.js,… | 93 | 1218 | active |
| gali8/Tesseract-OCR-iOS An iOS framework wrapping the Tesseract OCR engine (with Leptonica and image libraries) for use in Objective-C or Swift apps on iOS 9.0+. I… | 23 | 4221 | maintenance |
| frotms/PaddleOCR2Pytorch A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the… | 73 | 1205 | active |
| opensemanticsearch/open-semantic-search An open-source integrated search server and ETL framework for processing, analyzing, and exploring large document collections. It combines … | 40 | 1203 | active |
| Calamari-OCR/calamari Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning… | 74 | 1197 | active |
| lessthanoptimal/BoofCV BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur… | 86 | 1192 | active |
| BnanZ0/ok-nte ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated… | 78 | 1174 | active |
| magicrew/doc7 doc7 is a Go CLI tool that converts PDFs, Office files, scans, screenshots, charts, formulas, and diagrams into AI-ready Markdown using any… | 77 | 1173 | active |
| jenly1314/MLKit MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin… | 79 | 1168 | active |
| DavidVentura/offline-translator An Android app that translates text, PDF/ODT documents, and images entirely offline using Firefox translation models on-device. It also off… | 88 | 1166 | active |
| CH563/shot-easy-website ShotEasy is a free online photo and screenshot toolkit built with Astro that runs entirely in the browser using WebAssembly. It offers scre… | 70 | 1162 | active |
| Open Food Facts Smooth App is the official Open Food Facts mobile application for Android and iOS, built with Flutter and Dart. It lets users scan food pro… | 95 | 1147 | active |
| ttttccxxui/DataInfra-RedactionEverything A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and pla… | 60 | 1147 | active |
| sbs20/scanservjs scanservjs is a self-hosted web UI frontend for SANE-compatible scanners, letting you share one or more scanners over a network from a Linu… | 95 | 1145 | active |
| gsidhu/buzee-tauri Buzee is a superfast full-text search application for Mac and Windows built with Tauri, Rust, and Svelte. It indexes local documents, image… | 52 | 1143 | active |
| suifengqjn/videoWater AI快剪 (videoWater) is a desktop application for fully automated batch video editing, built in Go. It bundles clipping, merging, watermarking… | 32 | 1143 | active |
| YuehaiTeam/cocogoat A browser-based toolbox for Genshin Impact that performs local achievement recognition using PaddleOCR and onnxruntime, plus achievement ma… | 76 | 1139 | active |
| clovaai/deep-text-recognition-benchmark Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio… | 32 | 3942 | maintenance |
| Udayraj123/OMRChecker OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It… | 67 | 1134 | active |
| sml2h3/ddddocr-fastapi A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc… | 23 | 1126 | active |
| XieZhiFa/IdCardOCR An Android OCR library for offline recognition of Chinese second-generation ID cards, driver's licenses, and passports. It extracts all fie… | 75 | 1125 | active |
| easydoc-ai/easydoc EasyDoc is a multimodal document processing API that converts unstructured documents like PDFs into hierarchical, machine-readable JSON. It… | 35 | 1115 | active |
| Anionex/agent-vision-toolkit A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, … | 79 | 1108 | active |
| pencilresearch/OpenScanner Open Scanner is a free, open-source document scanning app for iPhone built with Swift and SwiftUI. It captures receipts, notes, and documen… | 24 | 1104 | active |
| dominostars/playtranslate PlayTranslate is a real-time screen translation app for Android that captures game or app text via OCR and translates it, with support for … | 81 | 1102 | active |
| LoredCast/filewizard A self-hosted, browser-based web UI for converting files between many formats, running OCR on PDFs and images, transcribing audio with Whis… | 46 | 1102 | active |
| zclucas/RMT RMT (RuoMengTu) is a free, open-source macro and desktop automation tool built on AutoHotkey v2. It supports recording and playing keyboard… | 88 | 1101 | active |
| m417z/Textify Textify is a small Windows utility that lets users copy text from dialog boxes and controls that don't normally allow text selection. It wo… | 67 | 1097 | active |
| Bogdanovich77/DeekSeek-OCR---Dockerized-API A Dockerized REST API and batch processing scripts that convert PDF documents to Markdown using the DeepSeek-OCR model behind a FastAPI bac… | 38 | 1091 | active |
| mittagessen/kraken kraken is a turn-key OCR/HTR engine built on neural networks, optimized for historical and non-Latin script material. It provides trainable… | 99 | 1061 | active |
| smxiazi/NEW_xp_CAPTCHA xp_CAPTCHA is a Burp Suite extension (Java plugin) that automatically recognizes CAPTCHAs during brute-force attacks, using a companion Pyt… | 23 | 1050 | active |
| aiptimizer/TurboOCR TurboOCR is an extremely fast GPU-accelerated document parser written in C++ that combines OCR, layout analysis, table extraction, and form… | 82 | 1043 | active |
| chrisgrieser/shimmering-obsidian Shimmering Obsidian is an Alfred workflow providing dozens of features for controlling an Obsidian vault from the macOS launcher, including… | 77 | 1034 | active |
| antimatter15/ocrad.js Ocrad.js is a pure-JavaScript port of the Ocrad OCR engine, compiled to JavaScript via Emscripten, that converts scanned images of text bac… | 32 | 3517 | maintenance |
| spatie/pdf-to-text A PHP library that wraps the pdftotext binary to extract plain text from PDF files with a simple, fluent API. It requires the poppler-utils… | 67 | 1029 | stable |
| BlueArchiveArisHelper/BAAH BAAH (BlueArchive Aris Helper) is an open-source Python automation script with a GUI that automatically completes daily tasks in the mobile… | 90 | 1028 | active |
| Agentic Document Extraction (ADE) The official CLI for LandingAI's Agentic Document Extraction (ADE), which parses documents into grounded Markdown and elements and extracts… | 84 | 1028 | active |
| ocropus-archive/DUP-ocropy OCRopy is a collection of Python-based tools for document analysis and OCR, covering binarization, page layout analysis, and text line reco… | 10 | 3465 | maintenance |
| SnapXL/SnapX SnapX is a free, open-source, cross-platform screenshot and screen recording tool forked from ShareX, built with C# and Avalonia. It lets u… | 76 | 1018 | active |
| eragonruan/text-detection-ctpn A TensorFlow implementation of the Connectionist Text Proposal Network (CTPN) for detecting horizontal scene text in images. It includes pr… | 23 | 3429 | maintenance |
| clovaai/CRAFT-pytorch Official PyTorch implementation of CRAFT (Character Region Awareness for Text Detection), a scene text detector that localizes text by pred… | 32 | 3398 | maintenance |
| xiaofengShi/CHINESE-OCR An end-to-end Chinese scene-text OCR pipeline combining CTPN for text detection, a VGG16-based orientation classifier, and CRNN with CTC fo… | 76 | 2955 | maintenance |
| nickliqian/cnn_captcha A Python project that uses convolutional neural networks built with TensorFlow to recognize character-based image captchas. It packages val… | 32 | 2881 | maintenance |
| alisen39/TrWebOCR TrWebOCR is an open-source offline Chinese OCR service built on the Tr project, exposing both a web UI and HTTP API for text recognition. I… | 23 | 2878 | maintenance |
| YCG09/chinese_ocr An end-to-end Chinese OCR system implemented with TensorFlow and Keras, combining CTPN for text detection with DenseNet + CTC for text reco… | 32 | 2782 | maintenance |
| ZBar/ZBar ZBar is an open-source C library and software suite for reading bar codes from video streams, image files, and raw intensity sensors. It su… | 32 | 2544 | maintenance |
| meijieru/crnn.pytorch A PyTorch implementation of the Convolutional Recurrent Neural Network (CRNN) for scene text recognition, based on the 2016 paper by Shi et… | 32 | 2492 | maintenance |
| ctripcorp/C-OCR C-OCR is Ctrip's in-house OCR project focused on recognizing travel-related documents such as ID cards, passports, train tickets, and visas… | 32 | 2476 | maintenance |
| alephdata/aleph Aleph is a self-hosted platform for indexing, searching, and browsing large volumes of documents (PDF, Word, HTML) and structured data (CSV… | 70 | 2420 | maintenance |
| Roujack/mathAI mathAI is a photo-based math problem solver written in Python: it takes an image containing a handwritten or printed arithmetic expression,… | 32 | 2370 | maintenance |
| MhLiao/DB A PyTorch implementation of DBNet and DBNet++, real-time arbitrary-shape scene text detection models based on differentiable binarization. … | 32 | 2260 | maintenance |
| githubharald/SimpleHTR A Handwritten Text Recognition (HTR) system implemented in TensorFlow that recognizes text from images of single words or text lines, train… | 72 | 2183 | maintenance |
| bgshih/crnn An implementation of the Convolutional Recurrent Neural Network (CRNN), combining CNN, RNN, and CTC loss for image-based sequence recogniti… | 32 | 2105 | maintenance |
| Pulover/PuloversMacroCreator Pulover's Macro Creator is a free Windows automation tool and script generator built on AutoHotkey, featuring a built-in recorder for keyst… | 23 | 2017 | maintenance |
| guanshuicheng/invoice A Flask-based OCR microservice that recognizes Chinese VAT invoices (electronic, regular, and special) using a YOLOv3 + CRNN + CTC deep lea… | 32 | 1982 | maintenance |
| PandaOCR PandaOCR is a free Windows desktop OCR tool that captures screen regions and recognizes text using many cloud OCR engines (Sogou, Tencent, … | 80 | 1918 | maintenance |
| Ucas-HaoranWei/Vary Official ECCV 2024 implementation of Vary, a method for scaling up the vision vocabulary of large vision-language models. It provides train… | 26 | 1889 | maintenance |
| Sierkinhane/CRNN_Chinese_Characters_Rec A PyTorch implementation of a CRNN (convolutional recurrent neural network) model for recognizing Chinese characters in images. It includes… | 32 | 1875 | maintenance |
| impira/docquery DocQuery is a Python library and CLI tool that uses large language models to answer questions about semi-structured and unstructured docume… | 32 | 1775 | maintenance |
| sergiomsilva/alpr-unconstrained An implementation of the ECCV 2018 paper 'License Plate Detection and Recognition in Unconstrained Scenarios', combining a Darknet-based de… | 32 | 1770 | maintenance |
| reworkd/tarsier Tarsier is a Python library providing vision utilities for LLM-driven web interaction agents. It visually tags interactable page elements w… | 17 | 1761 | maintenance |
| dbashford/textract A Node.js library that extracts plain text from many document formats including HTML, PDF, DOC/DOCX, XLS/XLSX, CSV, PPTX, RTF, EPUB, and im… | 58 | 1694 | maintenance |
| JonathanLink/PDFLayoutTextStripper A Java library that converts PDF files to text while preserving the original layout, built as a subclass of Apache PDFBox's PDFTextStripper… | 23 | 1608 | maintenance |
| DevashishPrasad/CascadeTabNet CascadeTabNet is a PyTorch/mmdetection implementation of a CVPR 2020 paper for end-to-end table detection and structure recognition from im… | 32 | 1549 | maintenance |
| allgood/OpenNoteScanner OpenNoteScanner is an open-source Android app for scanning handwritten notes and printed documents using a mobile device camera. It uses Op… | 31 | 1539 | maintenance |