function: ocr
484 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| MashiroSaber03/Saber-Translator Saber-Translator is an AI-powered manga translation application that detects speech bubbles, OCRs Japanese text, translates it, inpaints th… | 78 | 3516 | active |
| deepseek-ai/DeepSeek-OCR-2 DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn… | 44 | 3379 | active |
| InkTimeRecord/TTime TTime is a cross-platform desktop application for Windows and macOS that provides translation via input, screenshot, word selection, floati… | 22 | 3346 | active |
| deepdoctection/deepdoctection deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c… | 98 | 3248 | active |
| breezedeus/Pix2Text Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma… | 99 | 3227 | active |
| kerlomz/captcha_trainer A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren… | 55 | 3213 | active |
| CatchTheTornado/text-extract-api A self-hosted FastAPI-based API that converts PDFs, Office documents, and images into Markdown or structured JSON using OCR engines (EasyOC… | 45 | 3175 | active |
| Filimoa/open-parse Open Parse is a Python library that visually parses complex documents (primarily PDFs) into semantically meaningful chunks for LLM and RAG … | 64 | 3159 | active |
| otiai10/gosseract gosseract is a Go package that provides OCR (Optical Character Recognition) by binding to the Tesseract C++ library via cgo. It lets Go app… | 51 | 3130 | active |
| sw33tLie/macshot Macshot is a free, open-source, native macOS screenshot and screen recording tool built with Swift and AppKit. It offers region/window capt… | 74 | 3123 | active |
| apache/pdfbox Apache PDFBox is an open source Java library for working with PDF documents, supporting creation, manipulation, and content extraction. It … | 77 | 3105 | stable |
| AnyListen/tools-ocr Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model… | 23 | 3065 | active |
| open-rpa/openrpa OpenRPA is a free, open-source, enterprise-grade Robotic Process Automation (RPA) tool with a visual workflow designer for Windows. It can … | 69 | 3054 | active |
| thiagoalessio/tesseract-ocr-for-php A PHP wrapper library around the Tesseract OCR command-line binary, providing a fluent API for extracting text from images. It supports mul… | 61 | 3040 | stable |
| Dicklesworthstone/llm_aided_ocr A Python tool that converts scanned PDFs to text via Tesseract OCR, then uses LLMs (local or API-based like OpenAI/Anthropic) to correct OC… | 71 | 2993 | active |
| zcaceres/markdownify-mcp A Model Context Protocol (MCP) server that converts PDFs, images, audio, Office documents, and web content (including YouTube transcripts a… | 81 | 2983 | active |
| NVIDIA/NeMo-Retriever NVIDIA's scalable document content and metadata extraction library (also known as NVIDIA Ingest) that splits documents, classifies and extr… | 81 | 2970 | active |
| microsoft/table-transformer Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un… | 23 | 2939 | active |
| ciur/papermerge Papermerge is an open source document management system (DMS) designed for scanned documents and digital archives. It performs OCR on PDF, … | 59 | 2934 | active |
| openrecall/openrecall OpenRecall is an open-source, privacy-first digital memory tool that periodically takes screenshots of your screen, extracts text via local… | 43 | 2934 | active |
| ogkalu2/comic-translate An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language… | 90 | 2911 | active |
| duongductrong/Snapzy Snapzy is a free, open-source native macOS app for screenshots, screen recording, annotation, and OCR, built with SwiftUI, AppKit, and Scre… | 78 | 2870 | active |
| openalpr/openalpr OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete… | 23 | 11452 | maintenance |
| imanoop7/Ollama-OCR A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports… | 26 | 2780 | active |
| kha-white/manga-ocr Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to… | 90 | 2758 | stable |
| Audiveris/audiveris Audiveris is an open-source Optical Music Recognition (OMR) application that transcribes scanned sheet music images into symbolic music dat… | 97 | 2727 | active |
| dynobo/normcap NormCap is an OCR-powered screen-capture application that lets users select a region of the screen and extracts its text to the clipboard i… | 68 | 2695 | active |
| naiveHobo/InvoiceNet InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN… | 32 | 2694 | active |
| icereed/paperless-gpt A self-hosted companion application for paperless-ngx that uses LLMs and vision models to auto-generate document titles, tags, and dates, a… | 87 | 2651 | active |
| hgmzhn/manga-translator-ui A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,… | 80 | 2651 | active |
| sismics/docs Teedy (formerly Sismics Docs) is an open-source, lightweight document management system (DMS) for individuals and businesses. It offers OCR… | 67 | 2561 | active |
| UglyToad/PdfPig PdfPig is a C#/.NET library for reading and extracting text, images, annotations, forms, and metadata from PDF files, ported from Apache PD… | 91 | 2549 | active |
| chatdoc-com/OCRFlux OCRFlux is a Python toolkit built on a 3B-parameter vision-language model that converts PDFs and images into clean Markdown, handling compl… | 53 | 2533 | active |
| facebookresearch/nougat Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py… | 23 | 10063 | maintenance |
| openpaperwork/paperwork Paperwork is a personal document manager for Linux and Windows that scans, OCRs, indexes, and organizes paper documents. It provides keywor… | 10 | 2432 | active |
| schappim/macOCR macOCR is a macOS command-line tool that captures a screen region you select and runs OCR on it, copying the recognized text (or QR/barcode… | 88 | 2426 | active |
| ballerine-io/ballerine Ballerine is an open-source infrastructure and data orchestration platform for identity verification (KYC/KYB), fraud prevention, and merch… | 74 | 2426 | active |
| X-PLUG/mPLUG-DocOwl mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO… | 39 | 2411 | active |
| ossappscollective/OSS-DocumentScanner OSS Document Scanner is a free, open-source, privacy-focused mobile app for scanning documents with automatic edge detection, editing, OCR,… | 91 | 2385 | active |
| Cicada000/VV A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d… | 35 | 2375 | active |
| lyqht/mini-qr Mini QR is a web app (also installable as a PWA) for creating stylish, customizable QR codes and scanning QR codes via camera or image uplo… | 96 | 2373 | active |
| eikek/docspell Docspell is a self-hosted document management system (DMS) for organizing scanned papers, e-mails, and other files. It automates metadata t… | 67 | 2314 | active |
| Achno/gowall Gowall is a Go-based CLI tool that converts images (especially wallpapers) to custom color schemes and offers a broad suite of image proces… | 73 | 2301 | active |
| wxyhgk/retain-pdf RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/… | 77 | 2222 | active |
| KartikLabhshetwar/better-shot BetterShot is a native macOS app combining screenshots, screen recording, and a multi-clip video editor, built in Swift/SwiftUI as an open-… | 82 | 2206 | active |
| invoice-x/invoice2data A Python library and CLI tool that extracts structured data from PDF invoices using pluggable text-extraction backends (including OCR) and … | 97 | 2203 | active |
| MarkPDFdown/markpdfdown MarkPDFDown is a Python CLI tool that converts PDF documents and images into clean Markdown using multimodal large language models via Lite… | 60 | 2184 | active |
| TimmyOVO/deepseek-ocr.rs A Rust implementation of the DeepSeek-OCR inference stack with multiple OCR/VLM backends (DeepSeek-OCR, PaddleOCR-VL, DotsOCR), DSQ quantiz… | 60 | 2182 | active |
| ningzimu/image-to-editable-ppt-skill A Codex skill that converts slide images, PDFs, and image-based PPTX files into fully editable PowerPoint decks. It normalizes inputs into … | 76 | 2180 | active |
| sirfz/tesserocr A Python wrapper around the tesseract-ocr C++ API built with Cython for optical character recognition. It is Pillow-friendly, works with im… | 93 | 2171 | active |
| scambier/obsidian-omnisearch Omnisearch is an Obsidian community plugin providing instant, relevance-weighted full-text search across notes, PDFs, Office documents, and… | 98 | 2127 | active |
| NanoNets/docext docext is an on-premises document intelligence toolkit powered by vision-language models, offering OCR-free structured data extraction, PDF… | 50 | 2085 | active |
| DanBloomberg/leptonica Leptonica is an open-source C library providing a broad set of image processing and image analysis operations, with a focus on document ima… | 74 | 2074 | stable |
| manisandro/gImageReader gImageReader is a graphical GTK/Qt front-end to the tesseract-ocr engine for recognizing text in images, PDFs, scans, and screenshots. It s… | 51 | 1984 | active |
| crow-translate/crow-translate Crow Translate is a lightweight C++/Qt desktop translator that translates and speaks text using Google, Yandex, Bing, LibreTranslate, and L… | 10 | 1978 | active |
| KIYI671/AhabAssistantLimbusCompany AALC is a Windows desktop assistant for the game Limbus Company that automates repetitive gameplay tasks using image recognition and OCR. I… | 83 | 1971 | active |
| mosheng1/QuickClipboard QuickClipboard is a cross-platform clipboard enhancement tool (currently Windows and Android) built with Tauri 2, Rust, and React. It autom… | 85 | 1948 | active |
| f0ng/captcha-killer-modified A modified version of the captcha-killer Burp Suite extension that intercepts captcha images from HTTP responses and recognizes them using … | 41 | 1948 | active |
| clawsoftware/clawPDF clawPDF is an open-source virtual printer for Windows that converts printed output into PDF, PDF/A, OCR text, SVG, and various image format… | 23 | 1940 | active |
| LingyiChen-AI/JadeAI JadeAI is an open-source, AI-powered resume builder with drag-and-drop editing, 50+ professional templates, and real-time AI optimization i… | 82 | 1931 | active |
| Tencent-Hunyuan/HunyuanOCR HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model from Tencent, with a unified inference environment, llama.cpp PC-side … | 59 | 1930 | active |
| riddleling/iOS-OCR-Server An iOS app that turns an iPhone into a local OCR server using Apple's Vision Framework, exposing an HTTP API and web interface for image te… | 74 | 1927 | active |
| amebalabs/TRex TRex is a macOS menu bar application that extracts text from any visible screen area using OCR, copying it directly to the clipboard. It wo… | 84 | 1896 | active |
| rdumasia303/deepseek_ocr_app A self-hosted OCR web application combining a React frontend with a FastAPI backend, powered by the DeepSeek-OCR model. It processes images… | 50 | 1892 | active |
| robertknight/ocrs Ocrs is a Rust library and CLI tool for optical character recognition that extracts text from images such as scanned documents, photos, and… | 69 | 1875 | active |
| we0091234/Chinese_license_plate_detection_recognition A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports … | 71 | 1868 | active |
| jingsongliujing/OnnxOCR A lightweight multilingual OCR library rebuilt from PaddleOCR models to run on ONNXRuntime, removing the PaddlePaddle dependency for fast i… | 74 | 1860 | active |
| pmh1314520/WebRPA WebRPA is an open-source, no-code visual RPA tool for building automation workflows by dragging and connecting modules, covering web scrapi… | 82 | 1856 | active |
| simonw/tools A collection of miscellaneous HTML+JavaScript single-page tools hosted at tools.simonwillison.net, almost entirely generated with LLMs as a… | 69 | 1840 | active |
| clovaai/donut Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e… | 23 | 6919 | maintenance |
| AlibabaResearch/AdvancedLiterateMachinery A collection of original OCR and document understanding models, algorithms, and benchmarks from Alibaba's Tongyi Lab, including models like… | 55 | 1834 | active |
| camelot-dev/excalibur Excalibur is a self-hosted web interface for extracting tabular data from text-based PDFs, built on top of the Camelot PDF table extraction… | 71 | 1815 | active |
| sajari/docconv A Go library (with a companion docd CLI/HTTP service) that converts PDF, DOC, DOCX, XML, HTML, RTF, ODT, Pages documents and images into pl… | 23 | 1788 | active |
| xyTom/snippai Snippai is an AI-powered snipping tool that captures screenshots and uses AI to extract structured content such as LaTeX formulas, text, ta… | 89 | 1782 | active |
| chrisryugj/kordoc kordoc is a TypeScript CLI and MCP server that parses Korean document formats (HWP 3.x/5.x, HWPX, HWPML, PDF, XLS/XLSX, DOCX, and images wi… | 81 | 1776 | active |
| ianzhao/textshot TextShot is a Python command-line tool that lets you draw a rectangle over any screen region and copies the recognized text to your clipboa… | 32 | 1774 | active |
| yobix-ai/extractous Extractous is a fast Rust library for extracting content and metadata from unstructured documents like PDF, Word, Excel, HTML, CSV, and ema… | 24 | 1772 | active |
| puffinsoft/jscanify jscanify is an open-source pure JavaScript document scanning library powered by OpenCV.js. It detects and highlights documents in images an… | 73 | 1769 | active |
| nguyenq/tess4j Tess4J is a Java JNA wrapper for the Tesseract OCR API, enabling optical character recognition in Java applications. It supports TIFF, JPEG… | 91 | 1757 | stable |
| Feather-2/Burner-X Paper Burner X is a browser-based AI workstation for processing, translating, and analyzing academic documents like PDFs, DOCX, PPTX, and E… | 51 | 1750 | active |
| LLM Sherpa LLM Sherpa is a Python client library providing APIs for layout-aware PDF and document parsing to feed LLM applications, backed by the open… | 18 | 1749 | active |
| wm94i/Work-Review Work Review is a local-first desktop application that automatically tracks which apps you used, websites you visited, window titles, usage … | 82 | 1748 | active |
| kmonkeyhead/MORT MORT is a Windows desktop application that extracts on-screen text in real time using OCR and translates it via databases or machine transl… | 99 | 1728 | active |
| liuruoze/EasyPR EasyPR is an open-source C++ library built on OpenCV for recognizing Chinese license plates in unconstrained situations, outputting plate c… | 23 | 6429 | maintenance |
| cxOrz/chaoxing-signin A Node.js/TypeScript tool that automates sign-in for the Chaoxing (Superstar Learning) online course platform, supporting normal, photo, ge… | 10 | 1724 | active |
| kha-white/mokuro mokuro is a Python tool that performs text detection and OCR on Japanese manga pages and generates overlay files (.mokuro or HTML) enabling… | 86 | 1712 | active |
| hangone/WeBan WeBan is a Python-based automation tool that automatically completes courses and exams on the Weiban (安全微伴) university safety education pla… | 85 | 1712 | active |
| aisingapore/TagUI TagUI is a free, open-source robotic process automation (RPA) tool from AI Singapore that lets users write simple text flows to automate re… | 65 | 6325 | maintenance |
| cubewhy/skid-homework A browser-based, AI-powered homework solver built with Next.js that sends images and PDFs of homework problems to a Gemini or OpenAI-compat… | 61 | 1683 | active |
| chrismattmann/tika-python Tika-Python is a Python client library for Apache Tika's REST services, enabling document parsing, text extraction, and MIME detection nati… | 85 | 1666 | active |
| MgArcher/Text_select_captcha A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio… | 69 | 1656 | active |
| chineseocr A Python OCR toolkit that combines YOLO3-based text detection with CRNN/Dense recognition for Chinese and English text in natural scene ima… | 32 | 6123 | maintenance |
| RQLuo/MixTeX-Latex-OCR MixTeX is a multimodal OCR application that recognizes LaTeX formulas, tables, and mixed Chinese/English text from images, running entirely… | 22 | 1637 | active |
| aahnik/tgcf tgcf is a Python-based Telegram message forwarding automation tool that syncs messages between source and destination chats using either bo… | 23 | 1617 | active |
| Tsuk1ko/cq-picsearcher-bot A Node.js QQ bot that performs reverse image searches via saucenao, ascii2d, soutubot.moe, and trace.moe, connecting to any OneBot 11-compa… | 74 | 1596 | active |
| enoch3712/ExtractThinker ExtractThinker is a Python document intelligence library that uses LLMs to extract and classify structured data from documents like PDFs, i… | 45 | 1595 | active |
| kotaro-kinoshita/yomitoku YomiToku is an AI-powered document image analysis engine specialized for Japanese, providing full-text OCR, layout analysis, table structur… | 87 | 1579 | active |
| Layout-Parser/layout-parser LayoutParser is a Python toolkit for deep learning based document image analysis, offering unified APIs for layout detection models, layout… | 23 | 5774 | maintenance |
| hanmin0822/MisakaTranslator MisakaTranslator is a Windows desktop application that provides real-time machine translation for Galgames, text-based games, and manga. It… | 25 | 5746 | maintenance |
| IDEA-Research/Rex-Omni Rex-Omni is a 3B-parameter multimodal large language model that unifies object detection, OCR, pointing, keypoint detection, and visual pro… | 47 | 1561 | active |