Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: ocr

484 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
MashiroSaber03/Saber-Translator
Saber-Translator is an AI-powered manga translation application that detects speech bubbles, OCRs Japanese text, translates it, inpaints th…
783516active
deepseek-ai/DeepSeek-OCR-2
DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn…
443379active
InkTimeRecord/TTime
TTime is a cross-platform desktop application for Windows and macOS that provides translation via input, screenshot, word selection, floati…
223346active
deepdoctection/deepdoctection
deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c…
983248active
breezedeus/Pix2Text
Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma…
993227active
kerlomz/captcha_trainer
A deep learning training tool for image CAPTCHA recognition built on TensorFlow, using CNN/ResNet/DenseNet backbones with GRU/LSTM recurren…
553213active
CatchTheTornado/text-extract-api
A self-hosted FastAPI-based API that converts PDFs, Office documents, and images into Markdown or structured JSON using OCR engines (EasyOC…
453175active
Filimoa/open-parse
Open Parse is a Python library that visually parses complex documents (primarily PDFs) into semantically meaningful chunks for LLM and RAG …
643159active
otiai10/gosseract
gosseract is a Go package that provides OCR (Optical Character Recognition) by binding to the Tesseract C++ library via cgo. It lets Go app…
513130active
sw33tLie/macshot
Macshot is a free, open-source, native macOS screenshot and screen recording tool built with Swift and AppKit. It offers region/window capt…
743123active
apache/pdfbox
Apache PDFBox is an open source Java library for working with PDF documents, supporting creation, manipulation, and content extraction. It …
773105stable
AnyListen/tools-ocr
Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model…
233065active
open-rpa/openrpa
OpenRPA is a free, open-source, enterprise-grade Robotic Process Automation (RPA) tool with a visual workflow designer for Windows. It can …
693054active
thiagoalessio/tesseract-ocr-for-php
A PHP wrapper library around the Tesseract OCR command-line binary, providing a fluent API for extracting text from images. It supports mul…
613040stable
Dicklesworthstone/llm_aided_ocr
A Python tool that converts scanned PDFs to text via Tesseract OCR, then uses LLMs (local or API-based like OpenAI/Anthropic) to correct OC…
712993active
zcaceres/markdownify-mcp
A Model Context Protocol (MCP) server that converts PDFs, images, audio, Office documents, and web content (including YouTube transcripts a…
812983active
NVIDIA/NeMo-Retriever
NVIDIA's scalable document content and metadata extraction library (also known as NVIDIA Ingest) that splits documents, classifies and extr…
812970active
microsoft/table-transformer
Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un…
232939active
ciur/papermerge
Papermerge is an open source document management system (DMS) designed for scanned documents and digital archives. It performs OCR on PDF, …
592934active
openrecall/openrecall
OpenRecall is an open-source, privacy-first digital memory tool that periodically takes screenshots of your screen, extracts text via local…
432934active
ogkalu2/comic-translate
An AI-powered application and browser extension that automatically translates comics, manga, manhwa, webtoons, and BDs across many language…
902911active
duongductrong/Snapzy
Snapzy is a free, open-source native macOS app for screenshots, screen recording, annotation, and OCR, built with SwiftUI, AppKit, and Scre…
782870active
openalpr/openalpr
OpenALPR is an open-source Automatic License Plate Recognition (ALPR) library written in C++ that analyzes images and video streams to dete…
2311452maintenance
imanoop7/Ollama-OCR
A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports…
262780active
kha-white/manga-ocr
Manga OCR is a Python library providing optical character recognition for Japanese text, focused on Japanese manga. It uses a custom end-to…
902758stable
Audiveris/audiveris
Audiveris is an open-source Optical Music Recognition (OMR) application that transcribes scanned sheet music images into symbolic music dat…
972727active
dynobo/normcap
NormCap is an OCR-powered screen-capture application that lets users select a region of the screen and extracts its text to the clipboard i…
682695active
naiveHobo/InvoiceNet
InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN…
322694active
icereed/paperless-gpt
A self-hosted companion application for paperless-ngx that uses LLMs and vision models to auto-generate document titles, tags, and dates, a…
872651active
hgmzhn/manga-translator-ui
A desktop GUI application built on manga-image-translator that automatically translates text in manga/comic images across Japanese, Korean,…
802651active
sismics/docs
Teedy (formerly Sismics Docs) is an open-source, lightweight document management system (DMS) for individuals and businesses. It offers OCR…
672561active
UglyToad/PdfPig
PdfPig is a C#/.NET library for reading and extracting text, images, annotations, forms, and metadata from PDF files, ported from Apache PD…
912549active
chatdoc-com/OCRFlux
OCRFlux is a Python toolkit built on a 3B-parameter vision-language model that converts PDFs and images into clean Markdown, handling compl…
532533active
facebookresearch/nougat
Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py…
2310063maintenance
openpaperwork/paperwork
Paperwork is a personal document manager for Linux and Windows that scans, OCRs, indexes, and organizes paper documents. It provides keywor…
102432active
schappim/macOCR
macOCR is a macOS command-line tool that captures a screen region you select and runs OCR on it, copying the recognized text (or QR/barcode…
882426active
ballerine-io/ballerine
Ballerine is an open-source infrastructure and data orchestration platform for identity verification (KYC/KYB), fraud prevention, and merch…
742426active
X-PLUG/mPLUG-DocOwl
mPLUG-DocOwl is a family of multimodal large language models from Alibaba for OCR-free document understanding, including DocOwl1.5 and DocO…
392411active
ossappscollective/OSS-DocumentScanner
OSS Document Scanner is a free, open-source, privacy-focused mobile app for scanning documents with automatic edge detection, editing, OCR,…
912385active
Cicada000/VV
A Python application that indexes video clips of a public figure (mainly from the TV show 'This Is China') by combining face recognition (d…
352375active
lyqht/mini-qr
Mini QR is a web app (also installable as a PWA) for creating stylish, customizable QR codes and scanning QR codes via camera or image uplo…
962373active
eikek/docspell
Docspell is a self-hosted document management system (DMS) for organizing scanned papers, e-mails, and other files. It automates metadata t…
672314active
Achno/gowall
Gowall is a Go-based CLI tool that converts images (especially wallpapers) to custom color schemes and offers a broad suite of image proces…
732301active
wxyhgk/retain-pdf
RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/…
772222active
KartikLabhshetwar/better-shot
BetterShot is a native macOS app combining screenshots, screen recording, and a multi-clip video editor, built in Swift/SwiftUI as an open-…
822206active
invoice-x/invoice2data
A Python library and CLI tool that extracts structured data from PDF invoices using pluggable text-extraction backends (including OCR) and …
972203active
MarkPDFdown/markpdfdown
MarkPDFDown is a Python CLI tool that converts PDF documents and images into clean Markdown using multimodal large language models via Lite…
602184active
TimmyOVO/deepseek-ocr.rs
A Rust implementation of the DeepSeek-OCR inference stack with multiple OCR/VLM backends (DeepSeek-OCR, PaddleOCR-VL, DotsOCR), DSQ quantiz…
602182active
ningzimu/image-to-editable-ppt-skill
A Codex skill that converts slide images, PDFs, and image-based PPTX files into fully editable PowerPoint decks. It normalizes inputs into …
762180active
sirfz/tesserocr
A Python wrapper around the tesseract-ocr C++ API built with Cython for optical character recognition. It is Pillow-friendly, works with im…
932171active
scambier/obsidian-omnisearch
Omnisearch is an Obsidian community plugin providing instant, relevance-weighted full-text search across notes, PDFs, Office documents, and…
982127active
NanoNets/docext
docext is an on-premises document intelligence toolkit powered by vision-language models, offering OCR-free structured data extraction, PDF…
502085active
DanBloomberg/leptonica
Leptonica is an open-source C library providing a broad set of image processing and image analysis operations, with a focus on document ima…
742074stable
manisandro/gImageReader
gImageReader is a graphical GTK/Qt front-end to the tesseract-ocr engine for recognizing text in images, PDFs, scans, and screenshots. It s…
511984active
crow-translate/crow-translate
Crow Translate is a lightweight C++/Qt desktop translator that translates and speaks text using Google, Yandex, Bing, LibreTranslate, and L…
101978active
KIYI671/AhabAssistantLimbusCompany
AALC is a Windows desktop assistant for the game Limbus Company that automates repetitive gameplay tasks using image recognition and OCR. I…
831971active
mosheng1/QuickClipboard
QuickClipboard is a cross-platform clipboard enhancement tool (currently Windows and Android) built with Tauri 2, Rust, and React. It autom…
851948active
f0ng/captcha-killer-modified
A modified version of the captcha-killer Burp Suite extension that intercepts captcha images from HTTP responses and recognizes them using …
411948active
clawsoftware/clawPDF
clawPDF is an open-source virtual printer for Windows that converts printed output into PDF, PDF/A, OCR text, SVG, and various image format…
231940active
LingyiChen-AI/JadeAI
JadeAI is an open-source, AI-powered resume builder with drag-and-drop editing, 50+ professional templates, and real-time AI optimization i…
821931active
Tencent-Hunyuan/HunyuanOCR
HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model from Tencent, with a unified inference environment, llama.cpp PC-side …
591930active
riddleling/iOS-OCR-Server
An iOS app that turns an iPhone into a local OCR server using Apple's Vision Framework, exposing an HTTP API and web interface for image te…
741927active
amebalabs/TRex
TRex is a macOS menu bar application that extracts text from any visible screen area using OCR, copying it directly to the clipboard. It wo…
841896active
rdumasia303/deepseek_ocr_app
A self-hosted OCR web application combining a React frontend with a FastAPI backend, powered by the DeepSeek-OCR model. It processes images…
501892active
robertknight/ocrs
Ocrs is a Rust library and CLI tool for optical character recognition that extracts text from images such as scanned documents, photos, and…
691875active
we0091234/Chinese_license_plate_detection_recognition
A PyTorch-based Chinese license plate detection and recognition system built on YOLOv5 for detection and CRNN for recognition. It supports …
711868active
jingsongliujing/OnnxOCR
A lightweight multilingual OCR library rebuilt from PaddleOCR models to run on ONNXRuntime, removing the PaddlePaddle dependency for fast i…
741860active
pmh1314520/WebRPA
WebRPA is an open-source, no-code visual RPA tool for building automation workflows by dragging and connecting modules, covering web scrapi…
821856active
simonw/tools
A collection of miscellaneous HTML+JavaScript single-page tools hosted at tools.simonwillison.net, almost entirely generated with LLMs as a…
691840active
clovaai/donut
Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e…
236919maintenance
AlibabaResearch/AdvancedLiterateMachinery
A collection of original OCR and document understanding models, algorithms, and benchmarks from Alibaba's Tongyi Lab, including models like…
551834active
camelot-dev/excalibur
Excalibur is a self-hosted web interface for extracting tabular data from text-based PDFs, built on top of the Camelot PDF table extraction…
711815active
sajari/docconv
A Go library (with a companion docd CLI/HTTP service) that converts PDF, DOC, DOCX, XML, HTML, RTF, ODT, Pages documents and images into pl…
231788active
xyTom/snippai
Snippai is an AI-powered snipping tool that captures screenshots and uses AI to extract structured content such as LaTeX formulas, text, ta…
891782active
chrisryugj/kordoc
kordoc is a TypeScript CLI and MCP server that parses Korean document formats (HWP 3.x/5.x, HWPX, HWPML, PDF, XLS/XLSX, DOCX, and images wi…
811776active
ianzhao/textshot
TextShot is a Python command-line tool that lets you draw a rectangle over any screen region and copies the recognized text to your clipboa…
321774active
yobix-ai/extractous
Extractous is a fast Rust library for extracting content and metadata from unstructured documents like PDF, Word, Excel, HTML, CSV, and ema…
241772active
puffinsoft/jscanify
jscanify is an open-source pure JavaScript document scanning library powered by OpenCV.js. It detects and highlights documents in images an…
731769active
nguyenq/tess4j
Tess4J is a Java JNA wrapper for the Tesseract OCR API, enabling optical character recognition in Java applications. It supports TIFF, JPEG…
911757stable
Feather-2/Burner-X
Paper Burner X is a browser-based AI workstation for processing, translating, and analyzing academic documents like PDFs, DOCX, PPTX, and E…
511750active
LLM Sherpa
LLM Sherpa is a Python client library providing APIs for layout-aware PDF and document parsing to feed LLM applications, backed by the open…
181749active
wm94i/Work-Review
Work Review is a local-first desktop application that automatically tracks which apps you used, websites you visited, window titles, usage …
821748active
kmonkeyhead/MORT
MORT is a Windows desktop application that extracts on-screen text in real time using OCR and translates it via databases or machine transl…
991728active
liuruoze/EasyPR
EasyPR is an open-source C++ library built on OpenCV for recognizing Chinese license plates in unconstrained situations, outputting plate c…
236429maintenance
cxOrz/chaoxing-signin
A Node.js/TypeScript tool that automates sign-in for the Chaoxing (Superstar Learning) online course platform, supporting normal, photo, ge…
101724active
kha-white/mokuro
mokuro is a Python tool that performs text detection and OCR on Japanese manga pages and generates overlay files (.mokuro or HTML) enabling…
861712active
hangone/WeBan
WeBan is a Python-based automation tool that automatically completes courses and exams on the Weiban (安全微伴) university safety education pla…
851712active
aisingapore/TagUI
TagUI is a free, open-source robotic process automation (RPA) tool from AI Singapore that lets users write simple text flows to automate re…
656325maintenance
cubewhy/skid-homework
A browser-based, AI-powered homework solver built with Next.js that sends images and PDFs of homework problems to a Gemini or OpenAI-compat…
611683active
chrismattmann/tika-python
Tika-Python is a Python client library for Apache Tika's REST services, enabling document parsing, text extraction, and MIME detection nati…
851666active
MgArcher/Text_select_captcha
A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio…
691656active
chineseocr
A Python OCR toolkit that combines YOLO3-based text detection with CRNN/Dense recognition for Chinese and English text in natural scene ima…
326123maintenance
RQLuo/MixTeX-Latex-OCR
MixTeX is a multimodal OCR application that recognizes LaTeX formulas, tables, and mixed Chinese/English text from images, running entirely…
221637active
aahnik/tgcf
tgcf is a Python-based Telegram message forwarding automation tool that syncs messages between source and destination chats using either bo…
231617active
Tsuk1ko/cq-picsearcher-bot
A Node.js QQ bot that performs reverse image searches via saucenao, ascii2d, soutubot.moe, and trace.moe, connecting to any OneBot 11-compa…
741596active
enoch3712/ExtractThinker
ExtractThinker is a Python document intelligence library that uses LLMs to extract and classify structured data from documents like PDFs, i…
451595active
kotaro-kinoshita/yomitoku
YomiToku is an AI-powered document image analysis engine specialized for Japanese, providing full-text OCR, layout analysis, table structur…
871579active
Layout-Parser/layout-parser
LayoutParser is a Python toolkit for deep learning based document image analysis, offering unified APIs for layout detection models, layout…
235774maintenance
hanmin0822/MisakaTranslator
MisakaTranslator is a Windows desktop application that provides real-time machine translation for Galgames, text-based games, and manga. It…
255746maintenance
IDEA-Research/Rex-Omni
Rex-Omni is a 3B-parameter multimodal large language model that unifies object detection, OCR, pointing, keypoint detection, and visual pro…
471561active

← prev page 2 / 5 next →