Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: ocr

484 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
microsoft/markitdown
MarkItDown is a lightweight Python utility from Microsoft that converts many file formats (PDF, Office documents, images, audio, HTML, EPub…
83176487active
Stirling-Tools/Stirling-PDF
Stirling PDF is an open-source, self-hostable PDF platform offering 50-60+ tools for editing, merging, splitting, signing, redacting, conve…
9390501active
PaddlePaddle/PaddleOCR
PaddleOCR is a multilingual OCR and document parsing toolkit built on PaddlePaddle that converts images and PDFs into structured data like …
9388312stable
opendatalab/MinerU
MinerU is a document parsing tool that converts PDFs, images, DOCX, PPTX, and XLSX files into machine-readable Markdown and JSON. It handle…
8878560active
Tesseract OCR
Tesseract is an open-source OCR engine consisting of the libtesseract library and a command-line program, using an LSTM-based neural networ…
8676200stable
Docling
Docling is a Python library that parses and converts documents across many formats (PDF, DOCX, PPTX, XLSX, HTML, images, audio, and more) i…
8665603active
hiroi-sora/Umi-OCR
Umi-OCR is a free, open-source, fully offline OCR application for Windows and Linux with a Qt/QML GUI. It supports screenshot OCR, batch im…
4746882stable
paperless-ngx/paperless-ngx
Paperless-ngx is a self-hosted document management system that converts scanned physical documents into a searchable online archive. It use…
9544624active
ShareX/ShareX
ShareX is a free, open-source screen capture, screen recording, and file sharing application for Windows (with an Avalonia-based cross-plat…
9239322stable
datalab-to/marker
Marker is a Python library and CLI tool that converts PDFs, images, and office documents (DOCX, PPTX, XLSX, EPUB, HTML) into markdown, JSON…
9239295active
naptha/tesseract.js
Tesseract.js is a pure JavaScript port of the Tesseract OCR engine that extracts text from images in over 100 languages. It runs in the bro…
7038671active
ocrmypdf/OCRmyPDF
OCRmyPDF is a Python command-line tool that adds an OCR text layer to scanned PDF files using Tesseract, making them searchable and copy-pa…
9834589stable
JaidedAI/EasyOCR
EasyOCR is a ready-to-use Python OCR library built on PyTorch that extracts text from images, supporting 80+ languages and popular writing …
4829942stable
opendataloader-project/opendataloader-pdf
OpenDataLoader PDF is an open-source (Apache-2.0) PDF parser that converts PDFs into AI-ready Markdown, JSON with per-element bounding boxe…
8628817active
karakeep-app/karakeep
Karakeep (formerly Hoarder) is a self-hostable 'bookmark everything' app for saving links, notes, images, and PDFs with AI-based automatic …
9228617active
koodo-reader/koodo-reader
Koodo Reader is a cross-platform ebook manager and reader supporting EPUB, PDF, Kindle, comic archives, and many other formats. It offers c…
9427983active
microsoft/OmniParser
OmniParser is a screen parsing tool from Microsoft that converts UI screenshots into structured, understandable elements to ground vision-l…
6025310active
baidu/Unlimited-OCR
Baidu's Unlimited-OCR is an open vision-language OCR model for one-shot long-horizon document parsing, extending DeepSeek-OCR. It provides …
5624569active
deepseek-ai/DeepSeek-OCR
DeepSeek-OCR is an open vision-language model from DeepSeek AI that researches 'contexts optical compression' - encoding long text contexts…
4523855active
datalab-to/surya
Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, a…
8621318active
firecrawl/pdf-inspector
A fast Rust library for PDF inspection, classification, and text extraction that detects whether PDFs are text-based or scanned to enable s…
8116751active
lukas-blecher/LaTeX-OCR
pix2tex (LaTeX-OCR) is a PyTorch-based vision transformer model that converts images of math formulas into LaTeX code. It ships as a pip-in…
2416547stable
Unstructured-IO/unstructured
Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into…
9515349active
babalae/better-genshin-impact
BetterGI is a free, open-source Windows desktop application that automates gameplay in Genshin Impact using computer vision, OCR, and YOLO-…
9515067active
alam00000/bentopdf
BentoPDF is a self-hostable, privacy-first PDF toolkit that runs entirely client-side in the browser using WebAssembly, offering 50+ tools …
8214841active
ddddocr
DdddOcr is a Python library for offline, local recognition of various CAPTCHA types, including alphanumeric, Chinese character, and slider …
6414665active
tisfeng/Easydict
Easydict is a concise and elegant macOS dictionary and translation app for looking up words and translating text, ready to use out of the b…
9514372stable
T8RIN/ImageToolbox
Image Toolbox is a powerful open-source Android app for advanced image manipulation, built with Kotlin and Jetpack Compose in Material You …
9714364active
SubtitleEdit/subtitleedit
Subtitle Edit is a free, open-source desktop application for creating, editing, converting, and synchronizing subtitles, with video playbac…
9413971active
crimx/ext-saladict
Saladict is an open-source Chrome/Firefox WebExtension providing an all-in-one pop-up dictionary and page translator with multiple search m…
9613290active
HIllya51/LunaTranslator
LunaTranslator is a Windows application that translates visual novels (galgames) in real time. It extracts text via win32 hooking or OCR an…
9512912active
wmjordan/PDFPatcher
PDFPatcher is a free Windows PDF toolbox built on .NET with iText and MuPDF, offering bookmark editing, page cropping/rotation, merging and…
7012633active
DayBreak-u/chineseocr_lite
An ultra-lightweight Chinese OCR toolkit combining DBNet text detection, CRNN text recognition, and an angle classifier, with total model s…
7012339active
getomni-ai/zerox
Zerox is a library (Node.js and Python packages) that performs OCR and document extraction by converting files like PDFs, DOCX, and images …
3512266active
run-llama/liteparse
LiteParse is a fast, open-source document parser written in Rust that extracts spatial text with bounding boxes from PDFs, Office files, an…
7812181active
datalab-to/chandra
Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO…
7112171active
jsvine/pdfplumber
pdfplumber is a Python library for extracting detailed information from PDFs, including every character, line, rectangle, and table, built …
9010697active
PyMuPDF
PyMuPDF is a high-performance Python library built on the MuPDF C engine for extracting, analyzing, converting, rendering, and manipulating…
9810578stable
zyddnys/manga-image-translator
A Python tool that automatically detects, OCRs, translates, inpaints, and re-typesets text in images, primarily for manga and comics. It ru…
6510345active
CVHub520/X-AnyLabeling
X-AnyLabeling is a cross-platform desktop application for AI-assisted annotation of text, image, video, and multimodal data. It bundles bui…
9610212active
py-pdf/pypdf
pypdf is a free, open-source, pure-Python library for manipulating PDF files. It supports splitting, merging, cropping, and transforming pa…
9910173active
opendatalab/PDF-Extract-Kit
PDF-Extract-Kit is a Python model toolbox for high-quality PDF content extraction, integrating state-of-the-art models for layout detection…
259993active
ahrm/sioyek
Sioyek is a keyboard-focused PDF viewer designed for reading textbooks and research papers. It offers smart jumps to references, portals fo…
679805active
ripperhe/Bob
Bob is a macOS menu bar application for translation and OCR, supporting selection translation, screenshot translation, input translation, s…
499735active
YaoFANGUK/video-subtitle-extractor
A GUI application that extracts hard-coded (burned-in) subtitles from videos and generates SRT subtitle files using local deep-learning-bas…
709401active
studio-dots-ai/dots.ocr
dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi…
519090active
bytedance/Dolphin
Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-…
529049active
xitanggg/open-resume
OpenResume is an open-source web application that combines a resume builder and a resume parser. It generates modern, ATS-friendly resume P…
298862active
PantsuDango/Dango-Translator
Dango-Translator (团子翻译器) is a Windows desktop application that performs real-time OCR-based translation of on-screen text ('raw' untranslat…
918751active
Ucas-HaoranWei/GOT-OCR2.0
Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, …
258216active
STranslate/STranslate
STranslate is a ready-to-go Windows desktop translation and OCR tool built with WPF. It aggregates dozens of translation services (OpenAI, …
957860active
adithya-s-k/omniparse
OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into…
497815active
RapidAI/RapidOCR
RapidOCR is an open-source, multi-language OCR toolkit that performs text detection and recognition using models converted to run on ONNX R…
977599active
FluentRead/FluentRead
FluentRead is an open-source browser extension that provides bilingual webpage translation, instant selection translation, and image text t…
717551active
QuivrHQ/MegaParse
MegaParse is a Python library that parses PDFs, Word, PowerPoint, Excel, CSV, and text documents into LLM-friendly formats with a focus on …
297413active
zai-org/GLM-OCR
GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa…
657366active
testerSunshine/12306
A Python-based ticket-sniping assistant for China Railway's 12306 booking system that automates login, CAPTCHA recognition, and ticket purc…
1034086maintenance
xushengfeng/eSearch
eSearch is a cross-platform desktop application (Electron) combining screenshot capture, offline OCR based on PaddleOCR, screen search, tra…
987036active
OneDragon-Anything/ZenlessZoneZero-OneDragon
A Python-based automation assistant for the game Zenless Zone Zero that uses image recognition and OCR to fully automate daily tasks, dunge…
907029active
pdfminer/pdfminer.six
pdfminer.six is a community-maintained Python library for parsing and analyzing PDF documents, focused on extracting and analyzing text dat…
757018active
OLMo
olmOCR is an open toolkit from Ai2 that converts PDFs and image-based documents into clean, reading-order Markdown using a fine-tuned 7B vi…
526648active
Yuliang-Liu/MonkeyOCR
MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to…
606635active
steipete/summarize
Summarize is a Node.js CLI and Chrome/Firefox extension that extracts clean text from web pages, PDFs, YouTube videos, podcasts, and audio/…
826577active
madmaze/pytesseract
Python-tesseract is a Python wrapper for Google's Tesseract-OCR engine that recognizes and extracts text embedded in images. It supports al…
646383stable
mindee/doctr
docTR is a Python OCR library that extracts text from documents and images using a two-stage deep learning approach: text detection followe…
906315active
szad670401/HyperLPR
HyperLPR3 is a high-performance open-source framework for recognizing Chinese license plates, built with deep learning and available as a P…
276255active
oomol-lab/pdf-craft
pdf-craft is a Python library that converts PDF files into Markdown or EPUB, with a focus on scanned books and documents. It uses OCR (loca…
876226active
joeseesun/qiaomu-anything-to-notebooklm
A Claude Code Skill that ingests content from 15+ sources (WeChat articles, web pages, YouTube, PDFs, EPUB, Office docs, audio) and uploads…
625824active
freedomofpress/dangerzone
Dangerzone is a desktop application from Freedom of the Press Foundation that converts untrusted documents (PDFs, office files, images) int…
885715active
ramjke/Translumo
Translumo is a Windows desktop application that performs real-time screen translation by capturing on-screen text with OCR and translating …
725697active
pdf2htmlEX/pdf2htmlEX
pdf2htmlEX is a command-line tool that converts PDF files into HTML while preserving text, fonts, and formatting using modern web technolog…
375589active
mayswind/ezbookkeeping
ezBookkeeping is a lightweight, self-hosted personal finance and bookkeeping application built with Go and Vue.js. It supports transaction …
935471active
BIT-DataLab/Edit-Banana
Edit Banana is an open-source Python framework that converts static images and PDFs of diagrams, flowcharts, and charts into fully editable…
605469active
mayocream/koharu
Koharu is a local-first desktop application that automates manga translation using machine learning, combining text/bubble detection, OCR, …
825410active
papra-hq/papra
Papra is a minimalistic, self-hostable document management and archiving platform for long-term storage and retrieval of documents. It offe…
815247active
katanaml/sparrow
Sparrow is an open-source framework for structured data extraction from documents (PDFs, images) using ML, LLMs, and Vision LLMs, with sche…
955202active
dmMaze/BallonsTranslator
A desktop GUI application that uses deep learning to automatically translate comics and manga, combining text detection, OCR, inpainting, a…
995065active
TheJoeFin/Text-Grab
Text Grab is a Windows OCR utility that extracts text from anywhere on screen — screenshots, images, videos, PDFs, or app windows — entirel…
944978active
mg-chao/snow-apps
Snow Apps is a C++ open-source suite containing Snow Shot, a screenshot capture and annotation tool, and Snow Image Viewer. Snow Shot offer…
674926active
runhey/OnmyojiAutoScript
OnmyojiAutoScript (OAS) is a free, open-source automation script for the mobile game Onmyoji, built on the AzurLaneAutoScript framework. It…
654815active
open-mmlab/mmocr
MMOCR is OpenMMLab's PyTorch-based toolbox for text detection, recognition, and key information extraction. It provides a model zoo of OCR …
234752active
JabRef/jabref
JabRef is an open-source, cross-platform desktop application for managing BibTeX and BibLaTeX (.bib) reference libraries. It helps research…
864648active
LaoFeng-mouse/flyingmouse-format
FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop…
794601active
spipm/Depixelization_poc
Depix is a proof-of-concept tool that recovers plaintext from pixelized screenshots by matching pixelated blocks against a rendered font se…
104551active
Aidoku/Aidoku
Aidoku is a free, open-source manga reading application for iOS, iPadOS, and macOS with no ads. It supports local CBZ files, self-hosted me…
924495active
cyanfish/naps2
NAPS2 is a free, open-source document scanning application for Windows, Mac, and Linux that supports WIA, TWAIN, SANE, and ESCL scanners an…
974464active
pot-app/pot-desktop
Pot is a cross-platform desktop application for hotkey-based text translation, screenshot OCR, and text-to-speech, built with Tauri. It sup…
6219343maintenance
kevin2li/PDF-Guru
PDF Guru Anki is a cross-platform desktop and mobile application that combines a comprehensive PDF toolbox with deep Anki integration, conv…
304217active
LmeSzinc/StarRailCopilot
StarRailCopilot is a Python-based automation bot for the game Honkai: Star Rail, built on the next-generation Alas framework. It automates …
724195active
hcfyapp/crx-selection-translate
Huaci Fanyi (Selection Translate) is a browser extension for Chrome, Edge, and Firefox that translates selected text, full web pages, scree…
324139active
lumina-ai-inc/chunkr
Chunkr is an open-source document intelligence API that performs layout analysis, OCR, and semantic chunking to convert PDFs, presentations…
554137active
umlx5h/LLPlayer
LLPlayer is a Windows media player built for language learning, featuring dual subtitles, AI-generated subtitles via Whisper ASR, real-time…
784049active
apache/tika
Apache Tika is a Java toolkit that detects file types and extracts text and metadata from over a thousand file formats (PDF, Office documen…
774011stable
yuka-friends/Windrecorder
Windrecorder is a local-first screen recording and memory search app for Windows, an open-source alternative to Rewind.ai and Microsoft Rec…
473926active
camelot-dev/camelot
Camelot is a Python library for extracting tabular data from PDFs, offering five parsers including heuristic (lattice, stream), text-alignm…
943811active
MaaEnd/MaaEnd
MaaEnd is a vision-AI-powered automation assistant for the game 'Arknights: Endfield', built on MaaFramework. It captures the screen, recog…
953712active
liustack/modlens
ModLens is a vision plugin for DeepSeek Harness (dsh) and other text-only coding agents that converts pasted images into structured JSON ev…
783700active
Belval/TextRecognitionDataGenerator
A Python library and CLI tool (trdg) that generates synthetic text images for training OCR and text recognition models. It supports multipl…
233691stable
SteveTheKiller/KillerPDF
KillerPDF is a free, open-source (GPL-3.0) PDF editor for Windows built with C# and WPF, offering viewing, annotation, OCR, form filling, s…
813656active
CosmosShadow/gptpdf
A small Python library that parses PDF files into Markdown using a vision-capable LLM such as GPT-4o. It uses PyMuPDF to detect non-text ar…
313564active

page 1 / 5 next →