Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: pdf

514 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
microsoft/markitdown
MarkItDown is a lightweight Python utility from Microsoft that converts many file formats (PDF, Office documents, images, audio, HTML, EPub…
83176487active
Stirling-Tools/Stirling-PDF
Stirling PDF is an open-source, self-hostable PDF platform offering 50-60+ tools for editing, merging, splitting, signing, redacting, conve…
9390501active
PaddlePaddle/PaddleOCR
PaddleOCR is a multilingual OCR and document parsing toolkit built on PaddlePaddle that converts images and PDFs into structured data like …
9388312stable
opendatalab/MinerU
MinerU is a document parsing tool that converts PDFs, images, DOCX, PPTX, and XLSX files into machine-readable Markdown and JSON. It handle…
8878560active
Tesseract OCR
Tesseract is an open-source OCR engine consisting of the libtesseract library and a command-line program, using an LSTM-based neural networ…
8676200stable
Docling
Docling is a Python library that parses and converts documents across many formats (PDF, DOCX, PPTX, XLSX, HTML, images, audio, and more) i…
8665603active
PDF.js
PDF.js is Mozilla's general-purpose PDF viewer and rendering library built with HTML5 and JavaScript. It parses and renders PDF documents e…
9853786stable
hiroi-sora/Umi-OCR
Umi-OCR is a free, open-source, fully offline OCR application for Windows and Linux with a Qt/QML GUI. It supports screenshot OCR, batch im…
4746882stable
paperless-ngx/paperless-ngx
Paperless-ngx is a self-hosted document management system that converts scanned physical documents into a searchable online archive. It use…
9544624active
datalab-to/marker
Marker is a Python library and CLI tool that converts PDFs, images, and office documents (DOCX, PPTX, XLSX, EPUB, HTML) into markdown, JSON…
9239295active
PDFMathTranslate/PDFMathTranslate
PDFMathTranslate (pdf2zh) is a tool that translates scientific PDF documents while preserving the original layout, including formulas, char…
6936368active
VectifyAI/PageIndex
PageIndex is a Python SDK and framework for vectorless, reasoning-based RAG that replaces vector similarity search with a hierarchical tree…
8735333active
ocrmypdf/OCRmyPDF
OCRmyPDF is a Python command-line tool that adds an OCR text layer to scanned PDF files using Tesseract, making them searchable and copy-pa…
9834589stable
parallax/jsPDF
jsPDF is a JavaScript library for generating PDF documents entirely client-side in the browser, with Node.js support as well. It lets devel…
8931289stable
koreader/koreader
KOReader is an open-source document viewer application primarily aimed at e-ink e-readers, supporting formats like PDF, DjVu, EPUB, FB2, Mo…
8929289active
opendataloader-project/opendataloader-pdf
OpenDataLoader PDF is an open-source (Apache-2.0) PDF parser that converts PDFs into AI-ready Markdown, JSON with per-element bounding boxe…
8628817active
kovidgoyal/calibre
calibre is a free, open-source, cross-platform e-book management application written in Python. It can view, convert, edit, and catalog e-b…
9825734stable
baidu/Unlimited-OCR
Baidu's Unlimited-OCR is an open vision-language OCR model for one-shot long-horizon document parsing, extending DeepSeek-OCR. It provides …
5624569active
HKUDS/RAG-Anything
RAG-Anything is an all-in-one Python framework for multimodal Retrieval-Augmented Generation, built on LightRAG. It treats text, images, ta…
8023081active
datalab-to/surya
Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, a…
8621318active
firecrawl/anydoc
A fast Rust library that converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF documents into clean GitHub-Flavored Markd…
7918563active
docusealco/docuseal
DocuSeal is an open-source, self-hostable web application for creating, filling, and electronically signing PDF documents, positioned as a …
9118386active
sumatrapdfreader/sumatrapdf
SumatraPDF is a free, open-source, lightweight multi-format document reader for Windows supporting PDF, EPUB, MOBI, CBZ/CBR, DjVu, XPS, CHM…
9217401active
diegomura/react-pdf
A React renderer for creating PDF files declaratively using React components, running in both the browser and Node.js. It supports flexbox …
9516761active
firecrawl/pdf-inspector
A fast Rust library for PDF inspection, classification, and text extraction that detects whether PDFs are text-based or scanned to enable s…
8116751active
Unstructured-IO/unstructured
Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into…
9515349active
xournalpp/xournalpp
Xournal++ is a cross-platform handwriting notetaking application written in C++ with GTK3, supporting pressure-sensitive stylus input from …
9415286active
alam00000/bentopdf
BentoPDF is a self-hostable, privacy-first PDF toolkit that runs entirely client-side in the browser using WebAssembly, offering 50+ tools …
8214841active
kekingcn/kkFileView
kkFileView is a self-hosted Spring Boot application that provides online preview of a very wide range of file formats, including Office doc…
9514595active
QuestPDF/QuestPDF
QuestPDF is a modern C# library for generating PDF documents programmatically using a fluent, strongly-typed, component-based API. It inclu…
9914162stable
gotenberg/gotenberg
Gotenberg is a Docker-based HTTP API for converting documents (HTML, URLs, Markdown, Office files) into PDFs using headless Chromium and Li…
9812942stable
wmjordan/PDFPatcher
PDFPatcher is a free Windows PDF toolbox built on .NET with iText and MuPDF, offering bookmark editing, page cropping/rotation, merging and…
7012633active
bpampuch/pdfmake
pdfmake is a pure JavaScript library for generating PDF documents on both the client and server side, built on top of PDFKit. It uses a dec…
9012336stable
getomni-ai/zerox
Zerox is a library (Node.js and Python packages) that performs OCR and document extraction by converting files like PDFs, DOCX, and images …
3512266active
run-llama/liteparse
LiteParse is a fast, open-source document parser written in Rust that extracts spatial text with bounding boxes from PDFs, Office files, an…
7812181active
datalab-to/chandra
Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO…
7112171active
RelaxedJS/ReLaXed
ReLaXed is a Node.js command-line tool that generates PDF documents from HTML or Pug templates, styled with CSS/SCSS and rendered via headl…
5011798active
windingwind/zotero-pdf-translate
A Zotero plugin that translates text from PDFs, EPubs, webpages, metadata, annotations, and notes using 20+ translation services. It integr…
9511625active
Dompdf
Dompdf is an HTML to PDF converter implemented as a PHP library, functioning as a mostly CSS 2.1-compliant HTML layout and rendering engine…
9511177stable
wojtekmaj/react-pdf
A React component library for rendering and displaying PDF documents in web applications, built on top of Mozilla's PDF.js. It lets develop…
9411153stable
foliojs/pdfkit
PDFKit is a JavaScript library for programmatically generating PDF documents in Node.js and the browser. It offers a chainable, canvas-like…
9510698active
jsvine/pdfplumber
pdfplumber is a Python library for extracting detailed information from PDFs, including every character, line, rectangle, and table, built …
9010697active
PyMuPDF
PyMuPDF is a high-performance Python library built on the MuPDF C engine for extracting, analyzing, converting, rendering, and manipulating…
9810578stable
py-pdf/pypdf
pypdf is a free, open-source, pure-Python library for manipulating PDF files. It supports splitting, merging, cropping, and transforming pa…
9910173active
iib0011/omni-tools
OmniTools is a self-hosted web application bundling a large collection of browser-based utilities for images, video, audio, PDFs, text, dat…
6910088active
opendatalab/PDF-Extract-Kit
PDF-Extract-Kit is a Python model toolbox for high-quality PDF content extraction, integrating state-of-the-art models for layout detection…
259993active
phiresky/ripgrep-all
rga (ripgrep-all) is a command-line search tool that wraps ripgrep to search regex patterns inside PDFs, e-books, Office documents, SQLite …
649825active
ahrm/sioyek
Sioyek is a keyboard-focused PDF viewer designed for reading textbooks and research papers. It offers smart jumps to references, portals fo…
679805active
yihong0618/bilingual_book_maker
A Python CLI tool that uses AI translation services (ChatGPT, Claude, Gemini, DeepL, and others) to create bilingual versions of epub, txt,…
769609active
Kozea/WeasyPrint
WeasyPrint is a Python visual rendering engine for HTML and CSS that exports documents to PDF, with a CSS layout engine designed for pagina…
909531active
funstory-ai/BabelDOC
BabelDOC is a Python library and CLI tool for translating PDF scientific papers with bilingual comparison output. It uses LLM-based transla…
819408active
xberg-io/xberg
Xberg is a polyglot document intelligence framework with a Rust core that extracts text, metadata, tables, images, and structured data from…
869222active
Future-House/paper-qa
PaperQA2 is a Python library and CLI for high-accuracy retrieval augmented generation (RAG) over PDFs, text, Office documents, and source c…
939103active
studio-dots-ai/dots.ocr
dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi…
519090active
bytedance/Dolphin
Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-…
529049active
pdfcpu/pdfcpu
pdfcpu is a PDF processing library and command-line tool written in Go, supporting validation, optimization, encryption, signing, merging, …
948801active
DImuthuUpe/AndroidPdfViewer
An Android library for displaying PDF documents in apps, built on PdfiumAndroid for decoding. It supports animations, gestures, zoom, and d…
558470active
PDFCraftTool/pdfcraft
PDFCraft is a free, privacy-focused PDF toolkit with 90+ tools (merge, split, compress, convert, edit, secure) that runs entirely in the br…
758340active
Ucas-HaoranWei/GOT-OCR2.0
Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, …
258216active
freeok/so-novel
So Novel is a Java-based tool for extracting structured content from web pages and exporting it as EPUB, TXT, or PDF ebooks. It offers CLI,…
947943active
adithya-s-k/omniparse
OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into…
497815active
PHPOffice/PHPWord
PHPWord is a pure PHP library for reading and writing word processing documents in formats such as DOCX, ODT, RTF, DOC, HTML, and PDF. It l…
677585stable
QuivrHQ/MegaParse
MegaParse is a Python library that parses PDFs, Word, PowerPoint, Excel, CSV, and text documents into LLM-friendly formats with a focus on …
297413active
zai-org/GLM-OCR
GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa…
657366active
barryvdh/laravel-dompdf
A Laravel wrapper around the Dompdf HTML-to-PDF converter, exposing a facade and service provider for rendering Blade views or HTML strings…
877285stable
Wandmalfarbe/pandoc-latex-template
Eisvogel is a clean pandoc LaTeX template for converting markdown documents to PDF or LaTeX, including a beamer variant for slides. It is d…
887237stable
Zipstack/unstract
Unstract is an open-source, LLM-driven platform that extracts structured JSON data from unstructured documents such as PDFs, images, and sc…
887172active
google-research/arxiv-latex-cleaner
A Python command-line tool from Google Research that cleans LaTeX source code of academic papers for arXiv submission. It removes comments,…
807022active
pdfminer/pdfminer.six
pdfminer.six is a community-maintained Python library for parsing and analyzing PDF documents, focused on extracting and analyzing text dat…
757018active
OpenSignLabs/OpenSign
OpenSign is a free, open-source document e-signing application and DocuSign alternative built with React, Node.js, and MongoDB. It supports…
946919active
OLMo
olmOCR is an open toolkit from Ai2 that converts PDFs and image-based documents into clean, reading-order Markdown using a fine-tuned 7B vi…
526648active
Yuliang-Liu/MonkeyOCR
MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to…
606635active
happycola233/tchMaterial-parser
A cross-platform GUI tool that parses and downloads electronic textbook PDFs from China's National Smart Education Platform for Primary and…
956291active
oomol-lab/pdf-craft
pdf-craft is a Python library that converts PDF files into Markdown or EPUB, with a focus on scanned books and documents. It uses OCR (loca…
876226active
pdfarranger/pdfarranger
PDF Arranger is a small Python-GTK desktop application for merging, splitting, rotating, cropping, and rearranging PDF pages through an int…
855813active
guaguastandup/zotero-pdf2zh
A Zotero plugin that integrates PDF2zh and PDF2zh_next to translate PDF documents directly inside Zotero while preserving formulas and layo…
885760active
501351981/vue-office
A collection of Vue components for previewing office documents (docx, xlsx/xls, pdf, pptx) in the browser, supporting Vue 2/3 and non-Vue f…
535696active
pdf2htmlEX/pdf2htmlEX
pdf2htmlEX is a command-line tool that converts PDF files into HTML while preserving text, fonts, and formatting using modern web technolog…
375589active
ciromattia/kcc
KCC (Kindle Comic Converter) is a Python application with a Qt6 GUI and CLI that converts comics and manga from image folders, CBZ/EPUB arc…
985555active
qpdf/qpdf
qpdf is a command-line tool and C++ library that performs content-preserving transformations on PDF files, supporting linearization, encryp…
985350stable
papra-hq/papra
Papra is a minimalistic, self-hostable document management and archiving platform for long-term storage and retrieval of documents. It offe…
815247active
spatie/browsershot
A PHP library that converts web pages or raw HTML into images, PDFs, or rendered HTML strings using Puppeteer and headless Chrome. It suppo…
915241active
grobidOrg/grobid
GROBID is a Java machine learning library that extracts, parses, and re-structures raw documents such as PDFs into structured XML/TEI, focu…
875105active
eKoopmans/html2pdf.js
html2pdf.js is a JavaScript library that converts webpages or DOM elements into printable PDF files entirely in the browser, built on html2…
874913active
prawnpdf/prawn
Prawn is a pure Ruby library for programmatically generating PDF documents, offering vector drawing, text rendering, image embedding, and e…
584820stable
pdfme/pdfme
pdfme is an open-source TypeScript PDF generation library with a React-based WYSIWYG template designer, form, and viewer. It uses JSON temp…
904788active
danburzo/percollate
Percollate is a Node.js command-line tool that converts web pages into readable, well-formatted PDF, EPUB, HTML, or Markdown documents. It …
534667active
LaoFeng-mouse/flyingmouse-format
FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop…
794601active
torakiki/pdfsam
PDFsam Basic is a free, open-source desktop application for splitting, merging, rotating, mixing and extracting pages from PDF files. It is…
934551active
crabbly/Print.js
Print.js is a tiny JavaScript library that helps printing from the web, supporting printing of PDF, HTML, image, and JSON content directly …
654545stable
KnpLabs/snappy
Snappy is a PHP library that wraps the wkhtmltopdf and wkhtmltoimage binaries to generate PDFs, snapshots, or thumbnails from URLs or HTML …
924476stable
cyanfish/naps2
NAPS2 is a free, open-source document scanning application for Windows, Mac, and Linux that supports WIA, TWAIN, SANE, and ESCL scanners an…
974464active
embedpdf/embed-pdf-viewer
EmbedPDF is an open-source, framework-agnostic JavaScript PDF viewer that works with React, Vue, Svelte, Preact, and vanilla JS. It offers …
874426active
LibrePDF/OpenPDF
OpenPDF is an open-source Java library (LGPL/MPL licensed, forked from iText 4) for creating, editing, rendering, and encrypting PDF docume…
934359active
jonaslejon/malicious-pdf
A Python CLI tool that generates 67 malicious PDF test files embedding callbacks for SSRF, XSS, XXE, NTLM credential theft, and data exfilt…
864254active
kevin2li/PDF-Guru
PDF Guru Anki is a cross-platform desktop and mobile application that combines a comprehensive PDF toolbox with deep Anki integration, conv…
304217active
hcfyapp/crx-selection-translate
Huaci Fanyi (Selection Translate) is a browser extension for Chrome, Edge, and Firefox that translates selected text, full web pages, scree…
324139active
lumina-ai-inc/chunkr
Chunkr is an open-source document intelligence API that performs layout analysis, OCR, and semantic chunking to convert PDFs, presentations…
554137active
ethanwillis/zotero-scihub
A Zotero/Juris-M add-on that automatically downloads PDFs for library items with a DOI from Sci-Hub. It adds a context menu action and auto…
444080active
grimmory-tools/grimmory
Grimmory is a self-hosted digital library server for ebooks, comics, documents, and audiobooks, deployed via Docker with a web-based interf…
824070active

page 1 / 6 next →