function: ocr
484 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| microsoft/markitdown MarkItDown is a lightweight Python utility from Microsoft that converts many file formats (PDF, Office documents, images, audio, HTML, EPub… | 83 | 176487 | active |
| Stirling-Tools/Stirling-PDF Stirling PDF is an open-source, self-hostable PDF platform offering 50-60+ tools for editing, merging, splitting, signing, redacting, conve… | 93 | 90501 | active |
| PaddlePaddle/PaddleOCR PaddleOCR is a multilingual OCR and document parsing toolkit built on PaddlePaddle that converts images and PDFs into structured data like … | 93 | 88312 | stable |
| opendatalab/MinerU MinerU is a document parsing tool that converts PDFs, images, DOCX, PPTX, and XLSX files into machine-readable Markdown and JSON. It handle… | 88 | 78560 | active |
| Tesseract OCR Tesseract is an open-source OCR engine consisting of the libtesseract library and a command-line program, using an LSTM-based neural networ… | 86 | 76200 | stable |
| Docling Docling is a Python library that parses and converts documents across many formats (PDF, DOCX, PPTX, XLSX, HTML, images, audio, and more) i… | 86 | 65603 | active |
| hiroi-sora/Umi-OCR Umi-OCR is a free, open-source, fully offline OCR application for Windows and Linux with a Qt/QML GUI. It supports screenshot OCR, batch im… | 47 | 46882 | stable |
| paperless-ngx/paperless-ngx Paperless-ngx is a self-hosted document management system that converts scanned physical documents into a searchable online archive. It use… | 95 | 44624 | active |
| ShareX/ShareX ShareX is a free, open-source screen capture, screen recording, and file sharing application for Windows (with an Avalonia-based cross-plat… | 92 | 39322 | stable |
| datalab-to/marker Marker is a Python library and CLI tool that converts PDFs, images, and office documents (DOCX, PPTX, XLSX, EPUB, HTML) into markdown, JSON… | 92 | 39295 | active |
| naptha/tesseract.js Tesseract.js is a pure JavaScript port of the Tesseract OCR engine that extracts text from images in over 100 languages. It runs in the bro… | 70 | 38671 | active |
| ocrmypdf/OCRmyPDF OCRmyPDF is a Python command-line tool that adds an OCR text layer to scanned PDF files using Tesseract, making them searchable and copy-pa… | 98 | 34589 | stable |
| JaidedAI/EasyOCR EasyOCR is a ready-to-use Python OCR library built on PyTorch that extracts text from images, supporting 80+ languages and popular writing … | 48 | 29942 | stable |
| opendataloader-project/opendataloader-pdf OpenDataLoader PDF is an open-source (Apache-2.0) PDF parser that converts PDFs into AI-ready Markdown, JSON with per-element bounding boxe… | 86 | 28817 | active |
| karakeep-app/karakeep Karakeep (formerly Hoarder) is a self-hostable 'bookmark everything' app for saving links, notes, images, and PDFs with AI-based automatic … | 92 | 28617 | active |
| koodo-reader/koodo-reader Koodo Reader is a cross-platform ebook manager and reader supporting EPUB, PDF, Kindle, comic archives, and many other formats. It offers c… | 94 | 27983 | active |
| microsoft/OmniParser OmniParser is a screen parsing tool from Microsoft that converts UI screenshots into structured, understandable elements to ground vision-l… | 60 | 25310 | active |
| baidu/Unlimited-OCR Baidu's Unlimited-OCR is an open vision-language OCR model for one-shot long-horizon document parsing, extending DeepSeek-OCR. It provides … | 56 | 24569 | active |
| deepseek-ai/DeepSeek-OCR DeepSeek-OCR is an open vision-language model from DeepSeek AI that researches 'contexts optical compression' - encoding long text contexts… | 45 | 23855 | active |
| datalab-to/surya Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, a… | 86 | 21318 | active |
| firecrawl/pdf-inspector A fast Rust library for PDF inspection, classification, and text extraction that detects whether PDFs are text-based or scanned to enable s… | 81 | 16751 | active |
| lukas-blecher/LaTeX-OCR pix2tex (LaTeX-OCR) is a PyTorch-based vision transformer model that converts images of math formulas into LaTeX code. It ships as a pip-in… | 24 | 16547 | stable |
| Unstructured-IO/unstructured Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into… | 95 | 15349 | active |
| babalae/better-genshin-impact BetterGI is a free, open-source Windows desktop application that automates gameplay in Genshin Impact using computer vision, OCR, and YOLO-… | 95 | 15067 | active |
| alam00000/bentopdf BentoPDF is a self-hostable, privacy-first PDF toolkit that runs entirely client-side in the browser using WebAssembly, offering 50+ tools … | 82 | 14841 | active |
| ddddocr DdddOcr is a Python library for offline, local recognition of various CAPTCHA types, including alphanumeric, Chinese character, and slider … | 64 | 14665 | active |
| tisfeng/Easydict Easydict is a concise and elegant macOS dictionary and translation app for looking up words and translating text, ready to use out of the b… | 95 | 14372 | stable |
| T8RIN/ImageToolbox Image Toolbox is a powerful open-source Android app for advanced image manipulation, built with Kotlin and Jetpack Compose in Material You … | 97 | 14364 | active |
| SubtitleEdit/subtitleedit Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac… | 94 | 13971 | active |
| crimx/ext-saladict Saladict is an open-source Chrome/Firefox WebExtension providing an all-in-one pop-up dictionary and page translator with multiple search m… | 96 | 13290 | active |
| HIllya51/LunaTranslator LunaTranslator is a Windows application that translates visual novels (galgames) in real time. It extracts text via win32 hooking or OCR an… | 95 | 12912 | active |
| wmjordan/PDFPatcher PDFPatcher is a free Windows PDF toolbox built on .NET with iText and MuPDF, offering bookmark editing, page cropping/rotation, merging and… | 70 | 12633 | active |
| DayBreak-u/chineseocr_lite An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s… | 70 | 12339 | active |
| getomni-ai/zerox Zerox is a library (Node.js and Python packages) that performs OCR and document extraction by converting files like PDFs, DOCX, and images … | 35 | 12266 | active |
| run-llama/liteparse LiteParse is a fast, open-source document parser written in Rust that extracts spatial text with bounding boxes from PDFs, Office files, an… | 78 | 12181 | active |
| datalab-to/chandra Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO… | 71 | 12171 | active |
| jsvine/pdfplumber pdfplumber is a Python library for extracting detailed information from PDFs, including every character, line, rectangle, and table, built … | 90 | 10697 | active |
| PyMuPDF PyMuPDF is a high-performance Python library built on the MuPDF C engine for extracting, analyzing, converting, rendering, and manipulating… | 98 | 10578 | stable |
| zyddnys/manga-image-translator A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru… | 65 | 10345 | active |
| CVHub520/X-AnyLabeling X-AnyLabeling is a cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It bundles bui… | 96 | 10212 | active |
| py-pdf/pypdf pypdf is a free, open-source, pure-Python library for manipulating PDF files. It supports splitting, merging, cropping, and transforming pa… | 99 | 10173 | active |
| opendatalab/PDF-Extract-Kit PDF-Extract-Kit is a Python model toolbox for high-quality PDF content extraction, integrating state-of-the-art models for layout detection… | 25 | 9993 | active |
| ahrm/sioyek Sioyek is a keyboard-focused PDF viewer designed for reading textbooks and research papers. It offers smart jumps to references, portals fo… | 67 | 9805 | active |
| ripperhe/Bob Bob is a macOS menu bar application for translation and OCR, supporting selection translation, screenshot translation, input translation, s… | 49 | 9735 | active |
| YaoFANGUK/video-subtitle-extractor A GUI application that extracts hard-coded (burned-in) subtitles from videos and generates SRT subtitle files using local deep-learning-bas… | 70 | 9401 | active |
| studio-dots-ai/dots.ocr dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi… | 51 | 9090 | active |
| bytedance/Dolphin Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-… | 52 | 9049 | active |
| xitanggg/open-resume OpenResume is an open-source web application that combines a resume builder and a resume parser. It generates modern, ATS-friendly resume P… | 29 | 8862 | active |
| PantsuDango/Dango-Translator Dango-Translator (团子翻译器) is a Windows desktop application that performs real-time OCR-based translation of on-screen text ('raw' untranslat… | 91 | 8751 | active |
| Ucas-HaoranWei/GOT-OCR2.0 Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, … | 25 | 8216 | active |
| STranslate/STranslate STranslate is a ready-to-go Windows desktop translation and OCR tool built with WPF. It aggregates dozens of translation services (OpenAI, … | 95 | 7860 | active |
| adithya-s-k/omniparse OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into… | 49 | 7815 | active |
| RapidAI/RapidOCR RapidOCR is an open-source, multi-language OCR toolkit that performs text detection and recognition using models converted to run on ONNX R… | 97 | 7599 | active |
| FluentRead/FluentRead FluentRead is an open-source browser extension that provides bilingual webpage translation, instant selection translation, and image text t… | 71 | 7551 | active |
| QuivrHQ/MegaParse MegaParse is a Python library that parses PDFs, Word, PowerPoint, Excel, CSV, and text documents into LLM-friendly formats with a focus on … | 29 | 7413 | active |
| zai-org/GLM-OCR GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa… | 65 | 7366 | active |
| testerSunshine/12306 A Python-based ticket-sniping assistant for China Railway's 12306 booking system that automates login, CAPTCHA recognition, and ticket purc… | 10 | 34086 | maintenance |
| xushengfeng/eSearch eSearch is a cross-platform desktop application (Electron) combining screenshot capture, offline OCR based on PaddleOCR, screen search, tra… | 98 | 7036 | active |
| OneDragon-Anything/ZenlessZoneZero-OneDragon A Python-based automation assistant for the game Zenless Zone Zero that uses image recognition and OCR to fully automate daily tasks, dunge… | 90 | 7029 | active |
| pdfminer/pdfminer.six pdfminer.six is a community-maintained Python library for parsing and analyzing PDF documents, focused on extracting and analyzing text dat… | 75 | 7018 | active |
| OLMo olmOCR is an open toolkit from Ai2 that converts PDFs and image-based documents into clean, reading-order Markdown using a fine-tuned 7B vi… | 52 | 6648 | active |
| Yuliang-Liu/MonkeyOCR MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to… | 60 | 6635 | active |
| steipete/summarize Summarize is a Node.js CLI and Chrome/Firefox extension that extracts clean text from web pages, PDFs, YouTube videos, podcasts, and audio/… | 82 | 6577 | active |
| madmaze/pytesseract Python-tesseract is a Python wrapper for Google's Tesseract-OCR engine that recognizes and extracts text embedded in images. It supports al… | 64 | 6383 | stable |
| mindee/doctr docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe… | 90 | 6315 | active |
| szad670401/HyperLPR HyperLPR3 is a high-performance open-source framework for recognizing Chinese license plates, built with deep learning and available as a P… | 27 | 6255 | active |
| oomol-lab/pdf-craft pdf-craft is a Python library that converts PDF files into Markdown or EPUB, with a focus on scanned books and documents. It uses OCR (loca… | 87 | 6226 | active |
| joeseesun/qiaomu-anything-to-notebooklm A Claude Code Skill that ingests content from 15+ sources (WeChat articles, web pages, YouTube, PDFs, EPUB, Office docs, audio) and uploads… | 62 | 5824 | active |
| freedomofpress/dangerzone Dangerzone is a desktop application from Freedom of the Press Foundation that converts untrusted documents (PDFs, office files, images) int… | 88 | 5715 | active |
| ramjke/Translumo Translumo is a Windows desktop application that performs real-time screen translation by capturing on-screen text with OCR and translating … | 72 | 5697 | active |
| pdf2htmlEX/pdf2htmlEX pdf2htmlEX is a command-line tool that converts PDF files into HTML while preserving text, fonts, and formatting using modern web technolog… | 37 | 5589 | active |
| mayswind/ezbookkeeping ezBookkeeping is a lightweight, self-hosted personal finance and bookkeeping application built with Go and Vue.js. It supports transaction … | 93 | 5471 | active |
| BIT-DataLab/Edit-Banana Edit Banana is an open-source Python framework that converts static images and PDFs of diagrams, flowcharts, and charts into fully editable… | 60 | 5469 | active |
| mayocream/koharu Koharu is a local-first desktop application that automates manga translation using machine learning, combining text/bubble detection, OCR, … | 82 | 5410 | active |
| papra-hq/papra Papra is a minimalistic, self-hostable document management and archiving platform for long-term storage and retrieval of documents. It offe… | 81 | 5247 | active |
| katanaml/sparrow Sparrow is an open-source framework for structured data extraction from documents (PDFs, images) using ML, LLMs, and Vision LLMs, with sche… | 95 | 5202 | active |
| dmMaze/BallonsTranslator A desktop GUI application that uses deep learning to automatically translate comics and manga, combining text detection, OCR, inpainting, a… | 99 | 5065 | active |
| TheJoeFin/Text-Grab Text Grab is a Windows OCR utility that extracts text from anywhere on screen — screenshots, images, videos, PDFs, or app windows — entirel… | 94 | 4978 | active |
| mg-chao/snow-apps Snow Apps is a C++ open-source suite containing Snow Shot, a screenshot capture and annotation tool, and Snow Image Viewer. Snow Shot offer… | 67 | 4926 | active |
| runhey/OnmyojiAutoScript OnmyojiAutoScript (OAS) is a free, open-source automation script for the mobile game Onmyoji, built on the AzurLaneAutoScript framework. It… | 65 | 4815 | active |
| open-mmlab/mmocr MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR … | 23 | 4752 | active |
| JabRef/jabref JabRef is an open-source, cross-platform desktop application for managing BibTeX and BibLaTeX (.bib) reference libraries. It helps research… | 86 | 4648 | active |
| LaoFeng-mouse/flyingmouse-format FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop… | 79 | 4601 | active |
| spipm/Depixelization_poc Depix is a proof-of-concept tool that recovers plaintext from pixelized screenshots by matching pixelated blocks against a rendered font se… | 10 | 4551 | active |
| Aidoku/Aidoku Aidoku is a free, open-source manga reading application for iOS, iPadOS, and macOS with no ads. It supports local CBZ files, self-hosted me… | 92 | 4495 | active |
| cyanfish/naps2 NAPS2 is a free, open-source document scanning application for Windows, Mac, and Linux that supports WIA, TWAIN, SANE, and ESCL scanners an… | 97 | 4464 | active |
| pot-app/pot-desktop Pot is a cross-platform desktop application for hotkey-based text translation, screenshot OCR, and text-to-speech, built with Tauri. It sup… | 62 | 19343 | maintenance |
| kevin2li/PDF-Guru PDF Guru Anki is a cross-platform desktop and mobile application that combines a comprehensive PDF toolbox with deep Anki integration, conv… | 30 | 4217 | active |
| LmeSzinc/StarRailCopilot StarRailCopilot is a Python-based automation bot for the game Honkai: Star Rail, built on the next-generation Alas framework. It automates … | 72 | 4195 | active |
| hcfyapp/crx-selection-translate Huaci Fanyi (Selection Translate) is a browser extension for Chrome, Edge, and Firefox that translates selected text, full web pages, scree… | 32 | 4139 | active |
| lumina-ai-inc/chunkr Chunkr is an open-source document intelligence API that performs layout analysis, OCR, and semantic chunking to convert PDFs, presentations… | 55 | 4137 | active |
| umlx5h/LLPlayer LLPlayer is a Windows media player built for language learning, featuring dual subtitles, AI-generated subtitles via Whisper ASR, real-time… | 78 | 4049 | active |
| apache/tika Apache Tika is a Java toolkit that detects file types and extracts text and metadata from over a thousand file formats (PDF, Office documen… | 77 | 4011 | stable |
| yuka-friends/Windrecorder Windrecorder is a local-first screen recording and memory search app for Windows, an open-source alternative to Rewind.ai and Microsoft Rec… | 47 | 3926 | active |
| camelot-dev/camelot Camelot is a Python library for extracting tabular data from PDFs, offering five parsers including heuristic (lattice, stream), text-alignm… | 94 | 3811 | active |
| MaaEnd/MaaEnd MaaEnd is a vision-AI-powered automation assistant for the game 'Arknights: Endfield', built on MaaFramework. It captures the screen, recog… | 95 | 3712 | active |
| liustack/modlens ModLens is a vision plugin for DeepSeek Harness (dsh) and other text-only coding agents that converts pasted images into structured JSON ev… | 78 | 3700 | active |
| Belval/TextRecognitionDataGenerator A Python library and CLI tool (trdg) that generates synthetic text images for training OCR and text recognition models. It supports multipl… | 23 | 3691 | stable |
| SteveTheKiller/KillerPDF KillerPDF is a free, open-source (GPL-3.0) PDF editor for Windows built with C# and WPF, offering viewing, annotation, OCR, form filling, s… | 81 | 3656 | active |
| CosmosShadow/gptpdf A small Python library that parses PDF files into Markdown using a vision-capable LLM such as GPT-4o. It uses PyMuPDF to detect non-text ar… | 31 | 3564 | active |
page 1 / 5 next →