Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: ocr

484 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
drmingler/docling-api
A self-hostable FastAPI backend service that converts documents (PDF, DOCX, PPTX, HTML, images, CSV, AsciiDoc, Markdown) into Markdown usin…
651557active
hiroi-sora/PaddleOCR-json
An offline OCR command-line executable compiled from PaddleOCR C++ that recognizes text in images and outputs results as JSON strings. It c…
301542active
studyhelperhelper/studyhelper
An Android app disguised as a Sudoku game that automates earning daily points in the Xuexi Qiangguo (学习强国) app. It uses accessibility servi…
231533active
NanoNets/docstrange
DocStrange is a Python library and tool that converts documents (PDF, DOCX, PPTX, XLSX, images, URLs) into Markdown, JSON, CSV, or HTML usi…
401531active
WenmuZhou/PytorchOCR
A PyTorch-based OCR toolkit that ports PaddleOCR models to PyTorch, supporting common text detection and recognition algorithms like the PP…
591523active
DocumindHQ/documind
Documind is an open-source Node.js library that uses AI/LLMs to extract structured JSON data from PDFs and other documents based on customi…
331521active
jonaswinkler/paperless-ng
Paperless-ng is a self-hosted document management system that scans, OCRs, indexes, and archives physical documents with full-text search a…
105414maintenance
Turing-Project/WriteGPT
WriteGPT is a generative text-creation AI framework built on GPT-2 and other models (EAST, CRNN, BERT), fine-tuned to generate Chinese exam…
235289maintenance
dadoonet/fscrawler
FSCrawler is a Java-based file system crawler that indexes binary documents (PDF, MS Office, Open Office) into Elasticsearch, tracking new,…
881450active
beyondtranslate/beyondtranslate-ce
BeyondTranslate (formerly Biyi) is a cross-platform desktop translation app for macOS, Windows, and Linux built with Flutter and Rust. It c…
671449active
serratus/quaggaJS
QuaggaJS is a barcode-scanner library written entirely in JavaScript that supports real-time localization and decoding of barcode types suc…
235207maintenance
sdcb/PaddleSharp
A .NET/C# wrapper around Baidu's PaddleInference C API, providing PaddleOCR, PaddleDetection, rotation detection, Chinese segmentation, and…
721441active
Topdu/OpenOCR
OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta…
581437active
blacklanternsecurity/MANSPIDER
MANSPIDER is a Python CLI tool that crawls SMB shares across entire networks to find files by filename or content, with regex support and t…
771406active
alibaba/Logics-Parsing
Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul…
541402active
sMythicalBird/ZenlessZoneZero-Auto
A Python-based automation framework for the game Zenless Zone Zero that uses image classification, template matching, and OCR to perform au…
221378active
zelon88/HRConvert2
HRConvert2 is a self-hosted, resource-aware file conversion server written in PHP that supports 488 file formats across documents, images, …
991364active
saucepleez/taskt
taskt (formerly sharpRPA) is a free, open-source robotic process automation (RPA) client built in C# on the .NET Framework. It provides a W…
381363active
huridocs/pdf-document-layout-analysis
A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl…
821346active
CoderWanFeng/python-office
python-office is a Python office-automation library that wraps PDF, Word, Excel, PPT, image, video, email, WeChat, and OCR operations behin…
671344active
wormtql/yas
Yas is a fast screen-scanning tool that uses a custom-trained SVTR OCR model to read Genshin Impact and Honkai: Star Rail artifact stats di…
401336active
GauravSingh9356/J.A.R.V.I.S
A Python-based voice-controlled personal assistant inspired by Iron Man's J.A.R.V.I.S. It combines speech recognition, text-to-speech, OCR,…
481332active
ChaokunHong/MetaScreener
MetaScreener is an open-source AI tool that automates title/abstract and full-text PDF screening for systematic reviews using an ensemble o…
711328active
deanmalmgren/textract
A Python library that extracts text from virtually any document format (PDF, DOCX, PPTX, HTML, images, and more) through a simple unified i…
934698maintenance
gavrielc/Nano-PDF
A Python CLI tool that edits PDF slides using natural language prompts, powered by Google's Gemini 3 Pro Image model. It renders pages to i…
411321active
CBIhalsen/PolyglotPDF
PolyglotPDF is a Python-based multilingual eBook and PDF translation tool that preserves original layouts while translating, supporting bot…
421315active
Intuition-Lab/personal-model
Personal Model (Persome) is a local-first macOS runtime that captures focused cross-app activity and builds an evidence-linked personal mem…
771314active
ndl-lab/ndlocr-lite
NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma…
771309active
ttop32/MouseTooltipTranslator
A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports …
951300active
wisupai/e2m
E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using…
231294active
sist2app/sist2
sist2 is a fast, multi-threaded file system indexer that extracts text, metadata, and thumbnails from common file types, with OCR support v…
991289active
lzhgus/Capso
Capso is a free, open-source native macOS app for screenshots and screen recording, built with Swift 6.0 and SwiftUI as an alternative to C…
811286active
jbarrow/commonforms
CommonForms is a Python package and CLI that uses trained object-detection models (FFDNet-S/L) to automatically detect form fields in a PDF…
651284active
flutter-ml/google_ml_kit_flutter
A set of Flutter plugins wrapping Google's ML Kit on-device machine learning APIs for Android and iOS. It provides vision APIs (barcode, fa…
761274active
Yuliang-Liu/MonkeyOCRv2
MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2…
581256active
unjs/unpdf
unpdf is a TypeScript library providing PDF extraction and rendering utilities that work across all JavaScript runtimes, including Node.js,…
931218active
gali8/Tesseract-OCR-iOS
An iOS framework wrapping the Tesseract OCR engine (with Leptonica and image libraries) for use in Objective-C or Swift apps on iOS 9.0+. I…
234221maintenance
frotms/PaddleOCR2Pytorch
A PyTorch port of PaddleOCR that lets you run PaddleOCR-trained models (detection, recognition, and document structure parsing) without the…
731205active
opensemanticsearch/open-semantic-search
An open-source integrated search server and ETL framework for processing, analyzing, and exploring large document collections. It combines …
401203active
Calamari-OCR/calamari
Calamari is a Python-based OCR engine for line-based automatic text recognition, built on OCRopy and Kraken with a TensorFlow deep-learning…
741197active
lessthanoptimal/BoofCV
BoofCV is an open-source, real-time computer vision library written entirely in Java, covering image processing, camera calibration, featur…
861192active
BnanZ0/ok-nte
ok-nte is a Windows automation tool for the game Neverness to Everness that uses screenshot recognition, OCR, audio feedback, and simulated…
781174active
magicrew/doc7
doc7 is a Go CLI tool that converts PDFs, Office files, scans, screenshots, charts, formulas, and diagrams into AI-ready Markdown using any…
771173active
jenly1314/MLKit
MLKit is an easy-to-use Kotlin wrapper library around Google ML Kit for Android, exposing text recognition, barcode scanning, image labelin…
791168active
DavidVentura/offline-translator
An Android app that translates text, PDF/ODT documents, and images entirely offline using Firefox translation models on-device. It also off…
881166active
CH563/shot-easy-website
ShotEasy is a free online photo and screenshot toolkit built with Astro that runs entirely in the browser using WebAssembly. It offers scre…
701162active
Open Food Facts
Smooth App is the official Open Food Facts mobile application for Android and iOS, built with Flutter and Dart. It lets users scan food pro…
951147active
ttttccxxui/DataInfra-RedactionEverything
A local-first redaction workbench that detects and anonymizes sensitive information in documents, scanned PDFs, images, Word files, and pla…
601147active
sbs20/scanservjs
scanservjs is a self-hosted web UI frontend for SANE-compatible scanners, letting you share one or more scanners over a network from a Linu…
951145active
gsidhu/buzee-tauri
Buzee is a superfast full-text search application for Mac and Windows built with Tauri, Rust, and Svelte. It indexes local documents, image…
521143active
suifengqjn/videoWater
AI快剪 (videoWater) is a desktop application for fully automated batch video editing, built in Go. It bundles clipping, merging, watermarking…
321143active
YuehaiTeam/cocogoat
A browser-based toolbox for Genshin Impact that performs local achievement recognition using PaddleOCR and onnxruntime, plus achievement ma…
761139active
clovaai/deep-text-recognition-benchmark
Official PyTorch implementation of a four-stage scene text recognition (OCR) framework from an ICCV 2019 paper, with training and evaluatio…
323942maintenance
Udayraj123/OMRChecker
OMRChecker is a Python application that reads and evaluates OMR (Optical Mark Recognition) sheets scanned via a scanner or phone camera. It…
671134active
sml2h3/ddddocr-fastapi
A minimal FastAPI-based REST API service wrapping the DdddOcr OCR engine, exposing endpoints for image text recognition, slide captcha matc…
231126active
XieZhiFa/IdCardOCR
An Android OCR library for offline recognition of Chinese second-generation ID cards, driver's licenses, and passports. It extracts all fie…
751125active
easydoc-ai/easydoc
EasyDoc is a multimodal document processing API that converts unstructured documents like PDFs into hierarchical, machine-readable JSON. It…
351115active
Anionex/agent-vision-toolkit
A Python toolkit of vision CLI tools plus an agent skill that gives text-only LLM agents (like DeepSeek) image understanding capabilities, …
791108active
pencilresearch/OpenScanner
Open Scanner is a free, open-source document scanning app for iPhone built with Swift and SwiftUI. It captures receipts, notes, and documen…
241104active
dominostars/playtranslate
PlayTranslate is a real-time screen translation app for Android that captures game or app text via OCR and translates it, with support for …
811102active
LoredCast/filewizard
A self-hosted, browser-based web UI for converting files between many formats, running OCR on PDFs and images, transcribing audio with Whis…
461102active
zclucas/RMT
RMT (RuoMengTu) is a free, open-source macro and desktop automation tool built on AutoHotkey v2. It supports recording and playing keyboard…
881101active
m417z/Textify
Textify is a small Windows utility that lets users copy text from dialog boxes and controls that don't normally allow text selection. It wo…
671097active
Bogdanovich77/DeekSeek-OCR---Dockerized-API
A Dockerized REST API and batch processing scripts that convert PDF documents to Markdown using the DeepSeek-OCR model behind a FastAPI bac…
381091active
mittagessen/kraken
kraken is a turn-key OCR/HTR engine built on neural networks, optimized for historical and non-Latin script material. It provides trainable…
991061active
smxiazi/NEW_xp_CAPTCHA
xp_CAPTCHA is a Burp Suite extension (Java plugin) that automatically recognizes CAPTCHAs during brute-force attacks, using a companion Pyt…
231050active
aiptimizer/TurboOCR
TurboOCR is an extremely fast GPU-accelerated document parser written in C++ that combines OCR, layout analysis, table extraction, and form…
821043active
chrisgrieser/shimmering-obsidian
Shimmering Obsidian is an Alfred workflow providing dozens of features for controlling an Obsidian vault from the macOS launcher, including…
771034active
antimatter15/ocrad.js
Ocrad.js is a pure-JavaScript port of the Ocrad OCR engine, compiled to JavaScript via Emscripten, that converts scanned images of text bac…
323517maintenance
spatie/pdf-to-text
A PHP library that wraps the pdftotext binary to extract plain text from PDF files with a simple, fluent API. It requires the poppler-utils…
671029stable
BlueArchiveArisHelper/BAAH
BAAH (BlueArchive Aris Helper) is an open-source Python automation script with a GUI that automatically completes daily tasks in the mobile…
901028active
Agentic Document Extraction (ADE)
The official CLI for LandingAI's Agentic Document Extraction (ADE), which parses documents into grounded Markdown and elements and extracts…
841028active
ocropus-archive/DUP-ocropy
OCRopy is a collection of Python-based tools for document analysis and OCR, covering binarization, page layout analysis, and text line reco…
103465maintenance
SnapXL/SnapX
SnapX is a free, open-source, cross-platform screenshot and screen recording tool forked from ShareX, built with C# and Avalonia. It lets u…
761018active
eragonruan/text-detection-ctpn
A TensorFlow implementation of the Connectionist Text Proposal Network (CTPN) for detecting horizontal scene text in images. It includes pr…
233429maintenance
clovaai/CRAFT-pytorch
Official PyTorch implementation of CRAFT (Character Region Awareness for Text Detection), a scene text detector that localizes text by pred…
323398maintenance
xiaofengShi/CHINESE-OCR
An end-to-end Chinese scene-text OCR pipeline combining CTPN for text detection, a VGG16-based orientation classifier, and CRNN with CTC fo…
762955maintenance
nickliqian/cnn_captcha
A Python project that uses convolutional neural networks built with TensorFlow to recognize character-based image captchas. It packages val…
322881maintenance
alisen39/TrWebOCR
TrWebOCR is an open-source offline Chinese OCR service built on the Tr project, exposing both a web UI and HTTP API for text recognition. I…
232878maintenance
YCG09/chinese_ocr
An end-to-end Chinese OCR system implemented with TensorFlow and Keras, combining CTPN for text detection with DenseNet + CTC for text reco…
322782maintenance
ZBar/ZBar
ZBar is an open-source C library and software suite for reading bar codes from video streams, image files, and raw intensity sensors. It su…
322544maintenance
meijieru/crnn.pytorch
A PyTorch implementation of the Convolutional Recurrent Neural Network (CRNN) for scene text recognition, based on the 2016 paper by Shi et…
322492maintenance
ctripcorp/C-OCR
C-OCR is Ctrip's in-house OCR project focused on recognizing travel-related documents such as ID cards, passports, train tickets, and visas…
322476maintenance
alephdata/aleph
Aleph is a self-hosted platform for indexing, searching, and browsing large volumes of documents (PDF, Word, HTML) and structured data (CSV…
702420maintenance
Roujack/mathAI
mathAI is a photo-based math problem solver written in Python: it takes an image containing a handwritten or printed arithmetic expression,…
322370maintenance
MhLiao/DB
A PyTorch implementation of DBNet and DBNet++, real-time arbitrary-shape scene text detection models based on differentiable binarization. …
322260maintenance
githubharald/SimpleHTR
A Handwritten Text Recognition (HTR) system implemented in TensorFlow that recognizes text from images of single words or text lines, train…
722183maintenance
bgshih/crnn
An implementation of the Convolutional Recurrent Neural Network (CRNN), combining CNN, RNN, and CTC loss for image-based sequence recogniti…
322105maintenance
Pulover/PuloversMacroCreator
Pulover's Macro Creator is a free Windows automation tool and script generator built on AutoHotkey, featuring a built-in recorder for keyst…
232017maintenance
guanshuicheng/invoice
A Flask-based OCR microservice that recognizes Chinese VAT invoices (electronic, regular, and special) using a YOLOv3 + CRNN + CTC deep lea…
321982maintenance
PandaOCR
PandaOCR is a free Windows desktop OCR tool that captures screen regions and recognizes text using many cloud OCR engines (Sogou, Tencent, …
801918maintenance
Ucas-HaoranWei/Vary
Official ECCV 2024 implementation of Vary, a method for scaling up the vision vocabulary of large vision-language models. It provides train…
261889maintenance
Sierkinhane/CRNN_Chinese_Characters_Rec
A PyTorch implementation of a CRNN (convolutional recurrent neural network) model for recognizing Chinese characters in images. It includes…
321875maintenance
impira/docquery
DocQuery is a Python library and CLI tool that uses large language models to answer questions about semi-structured and unstructured docume…
321775maintenance
sergiomsilva/alpr-unconstrained
An implementation of the ECCV 2018 paper 'License Plate Detection and Recognition in Unconstrained Scenarios', combining a Darknet-based de…
321770maintenance
reworkd/tarsier
Tarsier is a Python library providing vision utilities for LLM-driven web interaction agents. It visually tags interactable page elements w…
171761maintenance
dbashford/textract
A Node.js library that extracts plain text from many document formats including HTML, PDF, DOC/DOCX, XLS/XLSX, CSV, PPTX, RTF, EPUB, and im…
581694maintenance
JonathanLink/PDFLayoutTextStripper
A Java library that converts PDF files to text while preserving the original layout, built as a subclass of Apache PDFBox's PDFTextStripper…
231608maintenance
DevashishPrasad/CascadeTabNet
CascadeTabNet is a PyTorch/mmdetection implementation of a CVPR 2020 paper for end-to-end table detection and structure recognition from im…
321549maintenance
allgood/OpenNoteScanner
OpenNoteScanner is an open-source Android app for scanning handwritten notes and printed documents using a mobile device camera. It uses Op…
311539maintenance

← prev page 3 / 5 next →