domain: pdf
514 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| microsoft/markitdown MarkItDown is a lightweight Python utility from Microsoft that converts many file formats (PDF, Office documents, images, audio, HTML, EPub… | 83 | 176487 | active |
| Stirling-Tools/Stirling-PDF Stirling PDF is an open-source, self-hostable PDF platform offering 50-60+ tools for editing, merging, splitting, signing, redacting, conve… | 93 | 90501 | active |
| PaddlePaddle/PaddleOCR PaddleOCR is a multilingual OCR and document parsing toolkit built on PaddlePaddle that converts images and PDFs into structured data like … | 93 | 88312 | stable |
| opendatalab/MinerU MinerU is a document parsing tool that converts PDFs, images, DOCX, PPTX, and XLSX files into machine-readable Markdown and JSON. It handle… | 88 | 78560 | active |
| Tesseract OCR Tesseract is an open-source OCR engine consisting of the libtesseract library and a command-line program, using an LSTM-based neural networ… | 86 | 76200 | stable |
| Docling Docling is a Python library that parses and converts documents across many formats (PDF, DOCX, PPTX, XLSX, HTML, images, audio, and more) i… | 86 | 65603 | active |
| PDF.js PDF.js is Mozilla's general-purpose PDF viewer and rendering library built with HTML5 and JavaScript. It parses and renders PDF documents e… | 98 | 53786 | stable |
| hiroi-sora/Umi-OCR Umi-OCR is a free, open-source, fully offline OCR application for Windows and Linux with a Qt/QML GUI. It supports screenshot OCR, batch im… | 47 | 46882 | stable |
| paperless-ngx/paperless-ngx Paperless-ngx is a self-hosted document management system that converts scanned physical documents into a searchable online archive. It use… | 95 | 44624 | active |
| datalab-to/marker Marker is a Python library and CLI tool that converts PDFs, images, and office documents (DOCX, PPTX, XLSX, EPUB, HTML) into markdown, JSON… | 92 | 39295 | active |
| PDFMathTranslate/PDFMathTranslate PDFMathTranslate (pdf2zh) is a tool that translates scientific PDF documents while preserving the original layout, including formulas, char… | 69 | 36368 | active |
| VectifyAI/PageIndex PageIndex is a Python SDK and framework for vectorless, reasoning-based RAG that replaces vector similarity search with a hierarchical tree… | 87 | 35333 | active |
| ocrmypdf/OCRmyPDF OCRmyPDF is a Python command-line tool that adds an OCR text layer to scanned PDF files using Tesseract, making them searchable and copy-pa… | 98 | 34589 | stable |
| parallax/jsPDF jsPDF is a JavaScript library for generating PDF documents entirely client-side in the browser, with Node.js support as well. It lets devel… | 89 | 31289 | stable |
| koreader/koreader KOReader is an open-source document viewer application primarily aimed at e-ink e-readers, supporting formats like PDF, DjVu, EPUB, FB2, Mo… | 89 | 29289 | active |
| opendataloader-project/opendataloader-pdf OpenDataLoader PDF is an open-source (Apache-2.0) PDF parser that converts PDFs into AI-ready Markdown, JSON with per-element bounding boxe… | 86 | 28817 | active |
| kovidgoyal/calibre calibre is a free, open-source, cross-platform e-book management application written in Python. It can view, convert, edit, and catalog e-b… | 98 | 25734 | stable |
| baidu/Unlimited-OCR Baidu's Unlimited-OCR is an open vision-language OCR model for one-shot long-horizon document parsing, extending DeepSeek-OCR. It provides … | 56 | 24569 | active |
| HKUDS/RAG-Anything RAG-Anything is an all-in-one Python framework for multimodal Retrieval-Augmented Generation, built on LightRAG. It treats text, images, ta… | 80 | 23081 | active |
| datalab-to/surya Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, a… | 86 | 21318 | active |
| firecrawl/anydoc A fast Rust library that converts Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF documents into clean GitHub-Flavored Markd… | 79 | 18563 | active |
| docusealco/docuseal DocuSeal is an open-source, self-hostable web application for creating, filling, and electronically signing PDF documents, positioned as a … | 91 | 18386 | active |
| sumatrapdfreader/sumatrapdf SumatraPDF is a free, open-source, lightweight multi-format document reader for Windows supporting PDF, EPUB, MOBI, CBZ/CBR, DjVu, XPS, CHM… | 92 | 17401 | active |
| diegomura/react-pdf A React renderer for creating PDF files declaratively using React components, running in both the browser and Node.js. It supports flexbox … | 95 | 16761 | active |
| firecrawl/pdf-inspector A fast Rust library for PDF inspection, classification, and text extraction that detects whether PDFs are text-based or scanned to enable s… | 81 | 16751 | active |
| Unstructured-IO/unstructured Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into… | 95 | 15349 | active |
| xournalpp/xournalpp Xournal++ is a cross-platform handwriting notetaking application written in C++ with GTK3, supporting pressure-sensitive stylus input from … | 94 | 15286 | active |
| alam00000/bentopdf BentoPDF is a self-hostable, privacy-first PDF toolkit that runs entirely client-side in the browser using WebAssembly, offering 50+ tools … | 82 | 14841 | active |
| kekingcn/kkFileView kkFileView is a self-hosted Spring Boot application that provides online preview of a very wide range of file formats, including Office doc… | 95 | 14595 | active |
| QuestPDF/QuestPDF QuestPDF is a modern C# library for generating PDF documents programmatically using a fluent, strongly-typed, component-based API. It inclu… | 99 | 14162 | stable |
| gotenberg/gotenberg Gotenberg is a Docker-based HTTP API for converting documents (HTML, URLs, Markdown, Office files) into PDFs using headless Chromium and Li… | 98 | 12942 | stable |
| wmjordan/PDFPatcher PDFPatcher is a free Windows PDF toolbox built on .NET with iText and MuPDF, offering bookmark editing, page cropping/rotation, merging and… | 70 | 12633 | active |
| bpampuch/pdfmake pdfmake is a pure JavaScript library for generating PDF documents on both the client and server side, built on top of PDFKit. It uses a dec… | 90 | 12336 | stable |
| getomni-ai/zerox Zerox is a library (Node.js and Python packages) that performs OCR and document extraction by converting files like PDFs, DOCX, and images … | 35 | 12266 | active |
| run-llama/liteparse LiteParse is a fast, open-source document parser written in Rust that extracts spatial text with bounding boxes from PDFs, Office files, an… | 78 | 12181 | active |
| datalab-to/chandra Chandra OCR 2 is a state-of-the-art open-weight OCR model from Datalab that converts images and PDFs into structured HTML, Markdown, or JSO… | 71 | 12171 | active |
| RelaxedJS/ReLaXed ReLaXed is a Node.js command-line tool that generates PDF documents from HTML or Pug templates, styled with CSS/SCSS and rendered via headl… | 50 | 11798 | active |
| windingwind/zotero-pdf-translate A Zotero plugin that translates text from PDFs, EPubs, webpages, metadata, annotations, and notes using 20+ translation services. It integr… | 95 | 11625 | active |
| Dompdf Dompdf is an HTML to PDF converter implemented as a PHP library, functioning as a mostly CSS 2.1-compliant HTML layout and rendering engine… | 95 | 11177 | stable |
| wojtekmaj/react-pdf A React component library for rendering and displaying PDF documents in web applications, built on top of Mozilla's PDF.js. It lets develop… | 94 | 11153 | stable |
| foliojs/pdfkit PDFKit is a JavaScript library for programmatically generating PDF documents in Node.js and the browser. It offers a chainable, canvas-like… | 95 | 10698 | active |
| jsvine/pdfplumber pdfplumber is a Python library for extracting detailed information from PDFs, including every character, line, rectangle, and table, built … | 90 | 10697 | active |
| PyMuPDF PyMuPDF is a high-performance Python library built on the MuPDF C engine for extracting, analyzing, converting, rendering, and manipulating… | 98 | 10578 | stable |
| py-pdf/pypdf pypdf is a free, open-source, pure-Python library for manipulating PDF files. It supports splitting, merging, cropping, and transforming pa… | 99 | 10173 | active |
| iib0011/omni-tools OmniTools is a self-hosted web application bundling a large collection of browser-based utilities for images, video, audio, PDFs, text, dat… | 69 | 10088 | active |
| opendatalab/PDF-Extract-Kit PDF-Extract-Kit is a Python model toolbox for high-quality PDF content extraction, integrating state-of-the-art models for layout detection… | 25 | 9993 | active |
| phiresky/ripgrep-all rga (ripgrep-all) is a command-line search tool that wraps ripgrep to search regex patterns inside PDFs, e-books, Office documents, SQLite … | 64 | 9825 | active |
| ahrm/sioyek Sioyek is a keyboard-focused PDF viewer designed for reading textbooks and research papers. It offers smart jumps to references, portals fo… | 67 | 9805 | active |
| yihong0618/bilingual_book_maker A Python CLI tool that uses AI translation services (ChatGPT, Claude, Gemini, DeepL, and others) to create bilingual versions of epub, txt,… | 76 | 9609 | active |
| Kozea/WeasyPrint WeasyPrint is a Python visual rendering engine for HTML and CSS that exports documents to PDF, with a CSS layout engine designed for pagina… | 90 | 9531 | active |
| funstory-ai/BabelDOC BabelDOC is a Python library and CLI tool for translating PDF scientific papers with bilingual comparison output. It uses LLM-based transla… | 81 | 9408 | active |
| xberg-io/xberg Xberg is a polyglot document intelligence framework with a Rust core that extracts text, metadata, tables, images, and structured data from… | 86 | 9222 | active |
| Future-House/paper-qa PaperQA2 is a Python library and CLI for high-accuracy retrieval augmented generation (RAG) over PDFs, text, Office documents, and source c… | 93 | 9103 | active |
| studio-dots-ai/dots.ocr dots.ocr is a 1.7B-parameter vision-language model for multilingual document layout parsing, converting documents into structured output wi… | 51 | 9090 | active |
| bytedance/Dolphin Dolphin is ByteDance's open-source document image parsing model that converts document images and PDFs into structured content using a two-… | 52 | 9049 | active |
| pdfcpu/pdfcpu pdfcpu is a PDF processing library and command-line tool written in Go, supporting validation, optimization, encryption, signing, merging, … | 94 | 8801 | active |
| DImuthuUpe/AndroidPdfViewer An Android library for displaying PDF documents in apps, built on PdfiumAndroid for decoding. It supports animations, gestures, zoom, and d… | 55 | 8470 | active |
| PDFCraftTool/pdfcraft PDFCraft is a free, privacy-focused PDF toolkit with 90+ tools (merge, split, compress, convert, edit, secure) that runs entirely in the br… | 75 | 8340 | active |
| Ucas-HaoranWei/GOT-OCR2.0 Official implementation of GOT-OCR2.0, a unified end-to-end vision-language model for general OCR that converts images of text, documents, … | 25 | 8216 | active |
| freeok/so-novel So Novel is a Java-based tool for extracting structured content from web pages and exporting it as EPUB, TXT, or PDF ebooks. It offers CLI,… | 94 | 7943 | active |
| adithya-s-k/omniparse OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into… | 49 | 7815 | active |
| PHPOffice/PHPWord PHPWord is a pure PHP library for reading and writing word processing documents in formats such as DOCX, ODT, RTF, DOC, HTML, and PDF. It l… | 67 | 7585 | stable |
| QuivrHQ/MegaParse MegaParse is a Python library that parses PDFs, Word, PowerPoint, Excel, CSV, and text documents into LLM-friendly formats with a focus on … | 29 | 7413 | active |
| zai-org/GLM-OCR GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understa… | 65 | 7366 | active |
| barryvdh/laravel-dompdf A Laravel wrapper around the Dompdf HTML-to-PDF converter, exposing a facade and service provider for rendering Blade views or HTML strings… | 87 | 7285 | stable |
| Wandmalfarbe/pandoc-latex-template Eisvogel is a clean pandoc LaTeX template for converting markdown documents to PDF or LaTeX, including a beamer variant for slides. It is d… | 88 | 7237 | stable |
| Zipstack/unstract Unstract is an open-source, LLM-driven platform that extracts structured JSON data from unstructured documents such as PDFs, images, and sc… | 88 | 7172 | active |
| google-research/arxiv-latex-cleaner A Python command-line tool from Google Research that cleans LaTeX source code of academic papers for arXiv submission. It removes comments,… | 80 | 7022 | active |
| pdfminer/pdfminer.six pdfminer.six is a community-maintained Python library for parsing and analyzing PDF documents, focused on extracting and analyzing text dat… | 75 | 7018 | active |
| OpenSignLabs/OpenSign OpenSign is a free, open-source document e-signing application and DocuSign alternative built with React, Node.js, and MongoDB. It supports… | 94 | 6919 | active |
| OLMo olmOCR is an open toolkit from Ai2 that converts PDFs and image-based documents into clean, reading-order Markdown using a fine-tuned 7B vi… | 52 | 6648 | active |
| Yuliang-Liu/MonkeyOCR MonkeyOCR is a lightweight large multimodal model (LMM) for document parsing that uses a Structure-Recognition-Relation triplet paradigm to… | 60 | 6635 | active |
| happycola233/tchMaterial-parser A cross-platform GUI tool that parses and downloads electronic textbook PDFs from China's National Smart Education Platform for Primary and… | 95 | 6291 | active |
| oomol-lab/pdf-craft pdf-craft is a Python library that converts PDF files into Markdown or EPUB, with a focus on scanned books and documents. It uses OCR (loca… | 87 | 6226 | active |
| pdfarranger/pdfarranger PDF Arranger is a small Python-GTK desktop application for merging, splitting, rotating, cropping, and rearranging PDF pages through an int… | 85 | 5813 | active |
| guaguastandup/zotero-pdf2zh A Zotero plugin that integrates PDF2zh and PDF2zh_next to translate PDF documents directly inside Zotero while preserving formulas and layo… | 88 | 5760 | active |
| 501351981/vue-office A collection of Vue components for previewing office documents (docx, xlsx/xls, pdf, pptx) in the browser, supporting Vue 2/3 and non-Vue f… | 53 | 5696 | active |
| pdf2htmlEX/pdf2htmlEX pdf2htmlEX is a command-line tool that converts PDF files into HTML while preserving text, fonts, and formatting using modern web technolog… | 37 | 5589 | active |
| ciromattia/kcc KCC (Kindle Comic Converter) is a Python application with a Qt6 GUI and CLI that converts comics and manga from image folders, CBZ/EPUB arc… | 98 | 5555 | active |
| qpdf/qpdf qpdf is a command-line tool and C++ library that performs content-preserving transformations on PDF files, supporting linearization, encryp… | 98 | 5350 | stable |
| papra-hq/papra Papra is a minimalistic, self-hostable document management and archiving platform for long-term storage and retrieval of documents. It offe… | 81 | 5247 | active |
| spatie/browsershot A PHP library that converts web pages or raw HTML into images, PDFs, or rendered HTML strings using Puppeteer and headless Chrome. It suppo… | 91 | 5241 | active |
| grobidOrg/grobid GROBID is a Java machine learning library that extracts, parses, and re-structures raw documents such as PDFs into structured XML/TEI, focu… | 87 | 5105 | active |
| eKoopmans/html2pdf.js html2pdf.js is a JavaScript library that converts webpages or DOM elements into printable PDF files entirely in the browser, built on html2… | 87 | 4913 | active |
| prawnpdf/prawn Prawn is a pure Ruby library for programmatically generating PDF documents, offering vector drawing, text rendering, image embedding, and e… | 58 | 4820 | stable |
| pdfme/pdfme pdfme is an open-source TypeScript PDF generation library with a React-based WYSIWYG template designer, form, and viewer. It uses JSON temp… | 90 | 4788 | active |
| danburzo/percollate Percollate is a Node.js command-line tool that converts web pages into readable, well-formatted PDF, EPUB, HTML, or Markdown documents. It … | 53 | 4667 | active |
| LaoFeng-mouse/flyingmouse-format FlyingMouse Format is an offline desktop file format converter for Windows (and macOS) built on Electron, bundling FFmpeg, LibreOffice, Pop… | 79 | 4601 | active |
| torakiki/pdfsam PDFsam Basic is a free, open-source desktop application for splitting, merging, rotating, mixing and extracting pages from PDF files. It is… | 93 | 4551 | active |
| crabbly/Print.js Print.js is a tiny JavaScript library that helps printing from the web, supporting printing of PDF, HTML, image, and JSON content directly … | 65 | 4545 | stable |
| KnpLabs/snappy Snappy is a PHP library that wraps the wkhtmltopdf and wkhtmltoimage binaries to generate PDFs, snapshots, or thumbnails from URLs or HTML … | 92 | 4476 | stable |
| cyanfish/naps2 NAPS2 is a free, open-source document scanning application for Windows, Mac, and Linux that supports WIA, TWAIN, SANE, and ESCL scanners an… | 97 | 4464 | active |
| embedpdf/embed-pdf-viewer EmbedPDF is an open-source, framework-agnostic JavaScript PDF viewer that works with React, Vue, Svelte, Preact, and vanilla JS. It offers … | 87 | 4426 | active |
| LibrePDF/OpenPDF OpenPDF is an open-source Java library (LGPL/MPL licensed, forked from iText 4) for creating, editing, rendering, and encrypting PDF docume… | 93 | 4359 | active |
| jonaslejon/malicious-pdf A Python CLI tool that generates 67 malicious PDF test files embedding callbacks for SSRF, XSS, XXE, NTLM credential theft, and data exfilt… | 86 | 4254 | active |
| kevin2li/PDF-Guru PDF Guru Anki is a cross-platform desktop and mobile application that combines a comprehensive PDF toolbox with deep Anki integration, conv… | 30 | 4217 | active |
| hcfyapp/crx-selection-translate Huaci Fanyi (Selection Translate) is a browser extension for Chrome, Edge, and Firefox that translates selected text, full web pages, scree… | 32 | 4139 | active |
| lumina-ai-inc/chunkr Chunkr is an open-source document intelligence API that performs layout analysis, OCR, and semantic chunking to convert PDFs, presentations… | 55 | 4137 | active |
| ethanwillis/zotero-scihub A Zotero/Juris-M add-on that automatically downloads PDFs for library items with a DOI from Sci-Hub. It adds a context menu action and auto… | 44 | 4080 | active |
| grimmory-tools/grimmory Grimmory is a self-hosted digital library server for ebooks, comics, documents, and audiobooks, deployed via Docker with a web-based interf… | 82 | 4070 | active |
page 1 / 6 next →