domain: pdf
514 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ofdrw/ofdrw OFDRW is an open-source Java library for reading, writing, and manipulating OFD (Open Fixed-layout Document) files, a Chinese national stan… | 93 | 1854 | active |
| mrmn2/PdfDing PdfDing is a self-hosted PDF manager, viewer, and editor with a browser-based interface that syncs reading progress across devices. It supp… | 10 | 1854 | active |
| clovaai/donut Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e… | 23 | 6919 | maintenance |
| realdennis/md2pdf An offline web application that converts Markdown files to PDF using a split-pane editor and browser print-to-PDF. It runs entirely client-… | 61 | 1829 | stable |
| jacoblee93/fully-local-pdf-chatbot A Next.js web application that lets users chat with uploaded PDF documents entirely locally, performing chunking, embedding, vector storage… | 52 | 1817 | active |
| nazdridoy/kokoro-tts A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w… | 88 | 1816 | active |
| camelot-dev/excalibur Excalibur is a self-hosted web interface for extracting tabular data from text-based PDFs, built on top of the Camelot PDF table extraction… | 71 | 1815 | active |
| wonday/react-native-pdf A React Native component that renders PDF documents on iOS and Android. It supports loading PDFs from URLs, blobs, local files, or assets w… | 90 | 1809 | active |
| LibPDF-js/core LibPDF is a modern TypeScript PDF library for parsing, modifying, generating, and signing PDF documents. It combines lenient parsing (like … | 80 | 1803 | active |
| sajari/docconv A Go library (with a companion docd CLI/HTTP service) that converts PDF, DOC, DOCX, XML, HTML, RTF, ODT, Pages documents and images into pl… | 23 | 1788 | active |
| sile-typesetter/sile SILE is a typesetting system that produces beautiful printed documents as PDF output, conceptually similar to TeX but written from scratch … | 76 | 1783 | active |
| sainnhe/caj2pdf-qt A cross-platform GUI application that converts CAJ, KDH, and NH files (CNKI academic document formats) to PDF, built with Qt on top of caj2… | 23 | 1780 | active |
| neuml/paperai paperai is an AI application for medical and scientific papers that runs bulk LLM inference and RAG pipelines over article repositories to … | 67 | 1779 | active |
| chrisryugj/kordoc kordoc is a TypeScript CLI and MCP server that parses Korean document formats (HWP 3.x/5.x, HWPX, HWPML, PDF, XLS/XLSX, DOCX, and images wi… | 81 | 1776 | active |
| yobix-ai/extractous Extractous is a fast Rust library for extracting content and metadata from unstructured documents like PDF, Word, Excel, HTML, CSV, and ema… | 24 | 1772 | active |
| spipu/html2pdf Html2Pdf is a PHP library that converts specially prepared HTML markup into PDF documents, built on top of TCPDF. It is designed for genera… | 41 | 1769 | active |
| golbin/hop HOP is an open-source desktop application for viewing and editing HWP/HWPX documents (the Korean Hangul word processor format) on macOS, Wi… | 77 | 1766 | active |
| nguyenq/tess4j Tess4J is a Java JNA wrapper for the Tesseract OCR API, enabling optical character recognition in Java applications. It supports TIFF, JPEG… | 91 | 1757 | stable |
| community-archive/obsidian-zotero-integration An Obsidian community plugin that inserts citations and bibliographies and imports notes and PDF annotations from Zotero into Obsidian. It … | 54 | 1757 | active |
| Feather-2/Burner-X Paper Burner X is a browser-based AI workstation for processing, translating, and analyzing academic documents like PDFs, DOCX, PPTX, and E… | 51 | 1750 | active |
| LLM Sherpa LLM Sherpa is a Python client library providing APIs for layout-aware PDF and document parsing to feed LLM applications, backed by the open… | 18 | 1749 | active |
| jesselau76/ebook-GPT-translator A Python CLI and GUI tool that translates ebooks (TXT, EPUB, DOCX, PDF, optional MOBI) using LLM providers like OpenAI, Azure, Codex, Claud… | 63 | 1728 | active |
| zstmfhy/zlibrary-to-notebooklm A Python CLI tool that automatically downloads books from Z-Library and uploads them to Google NotebookLM in one command. It uses Playwrigh… | 43 | 1711 | active |
| baskerville/plato Plato is a document reader application for Kobo e-ink e-readers, written in Rust. It supports PDF, EPUB, DJVU, CBZ, FB2, MOBI, XPS and TXT … | 65 | 1694 | active |
| pdf-rs/pdf A Rust library for reading, manipulating, and writing PDF files. Reading is fairly mature, while modifying and writing PDFs is still experi… | 74 | 1691 | active |
| foliojs/fontkit fontkit is an advanced font engine library for Node.js and the browser, used by PDFKit. It parses many font formats and provides glyph shap… | 23 | 1668 | stable |
| chrismattmann/tika-python Tika-Python is a Python client library for Apache Tika's REST services, enabling document parsing, text extraction, and MIME detection nati… | 85 | 1666 | active |
| Cimbali/pympress Pympress is a dual-screen PDF presentation tool that shows slides on a projector while giving the presenter a private view with notes, curr… | 56 | 1647 | active |
| lmn1919/dompdf.js dompdf.js is a pure-frontend DOM-to-PDF engine built with Rust, WebAssembly, and TypeScript that renders vector PDFs directly in the browse… | 79 | 1620 | active |
| mikehaertl/phpwkhtmltopdf A slim PHP wrapper around the wkhtmltopdf (and optionally wkhtmltoimage) command-line tools, exposing a clean OOP interface for creating PD… | 53 | 1608 | stable |
| clusterzx/paperless-ai A self-hosted web application that automatically analyzes, tags, and classifies documents in Paperless-ngx using LLMs via OpenAI-compatible… | 73 | 5910 | maintenance |
| LYiHub/mad-professor-public A desktop AI companion application for reading academic papers, built with PyQt6. It combines PDF parsing and translation, RAG-based retrie… | 27 | 1604 | active |
| enoch3712/ExtractThinker ExtractThinker is a Python document intelligence library that uses LLMs to extract and classify structured data from documents like PDFs, i… | 45 | 1595 | active |
| jodconverter/jodconverter JODConverter is a Java library that automates document conversions between office formats (e.g., docx, odt, xlsx, pdf) by driving LibreOffi… | 48 | 1594 | stable |
| kotaro-kinoshita/yomitoku YomiToku is an AI-powered document image analysis engine specialized for Japanese, providing full-text OCR, layout analysis, table structur… | 87 | 1579 | active |
| elementdavv/internet_archive_downloader A Chrome and Firefox browser extension that downloads borrowable books from Internet Archive (archive.org) and full-view books from HathiTr… | 75 | 1565 | active |
| drmingler/docling-api A self-hostable FastAPI backend service that converts documents (PDF, DOCX, PPTX, HTML, images, CSV, AsciiDoc, Markdown) into Markdown usin… | 65 | 1557 | active |
| LaravelDaily/laravel-invoices A Laravel PHP package for generating PDF invoice files from customizable parameters like items, taxes, discounts, and shipping. Invoices ca… | 68 | 1553 | active |
| alpha-liu-01/SpeedyNote SpeedyNote is a free, open-source, cross-platform note-taking application built in C++ with Qt, designed for stylus users who need fast PDF… | 87 | 1544 | active |
| 4lex4/scantailor-advanced ScanTailor Advanced is an interactive post-processing tool for scanned pages, merging features from ScanTailor Featured and Enhanced while … | 23 | 1532 | stable |
| NanoNets/docstrange DocStrange is a Python library and tool that converts documents (PDF, DOCX, PPTX, XLSX, images, URLs) into Markdown, JSON, CSV, or HTML usi… | 40 | 1531 | active |
| DavBfr/dart_pdf A Dart/Flutter library for creating PDF documents, offering both a low-level PDF generation API and a Flutter-like widget system for high-l… | 74 | 1527 | stable |
| kudrykv/latex-yearly-planner A Go/LaTeX tool that generates yearly PDF planners designed for e-ink tablets like Supernote and ReMarkable. It produces pre-generated plan… | 59 | 1525 | active |
| emcf/thepipe thepipe is a Python package and API that extracts clean markdown, multimodal media, and structured data from tricky documents like PDFs, Wo… | 58 | 1524 | active |
| oschwartz10612/poppler-windows A repository that packages prebuilt Poppler PDF binaries (from conda-forge's poppler-feedstock) with their dependencies into convenient Win… | 79 | 1522 | active |
| DocumindHQ/documind Documind is an open-source Node.js library that uses AI/LLMs to extract structured JSON data from PDFs and other documents based on customi… | 33 | 1521 | active |
| yamlresume/yamlresume YAMLResume is a resume-as-code tool that lets you write your CV in YAML and compile it to beautifully typeset PDFs via LaTeX. It ships as a… | 86 | 1504 | active |
| tjmlabs/ColiVara ColiVara is a hosted (and self-hostable) retrieval API that stores, searches, and retrieves documents using visual embeddings generated by … | 65 | 1486 | active |
| extend-hq/ui Extend UI is an open-source React component library (distributed via the shadcn registry) for building document-heavy product interfaces: P… | 58 | 1486 | active |
| KDE/okular Okular is KDE's universal document viewer for formats including PDF, PostScript, EPUB, Comic Book, and images, with support for annotations… | 77 | 1485 | stable |
| pagedjs/pagedjs Paged.js is an open-source JavaScript library that paginates HTML content in the browser, polyfilling the W3C CSS Paged Media and Generated… | 67 | 1485 | active |
| JakubMelka/PDF4QT PDF4QT is an open-source PDF editor, viewer, and rendering library written in C++ with Qt, implementing PDF 2.0 functionality. It ships a C… | 87 | 1459 | active |
| spatie/pdf-to-image A PHP library that converts PDF files to images (jpg, png, webp) using Imagick and Ghostscript. It supports rendering single or multiple pa… | 88 | 1457 | active |
| bblanchon/pdfium-binaries A project that automatically builds and distributes pre-compiled binaries of Google's PDFium PDF rendering library for many platforms and C… | 95 | 1456 | active |
| Open-Source-Legal/OpenContracts OpenContracts is a self-hosted, MIT-licensed document intelligence platform that turns document repositories into a programmable citation g… | 94 | 1451 | active |
| dadoonet/fscrawler FSCrawler is a Java-based file system crawler that indexes binary documents (PDF, MS Office, Open Office) into Elasticsearch, tracking new,… | 88 | 1450 | active |
| Nutlope/pdftochat PDFToChat is an open-source Next.js web application that lets users upload PDFs and chat with them using AI. It uses Together AI's Mixtral … | 68 | 1441 | active |
| Topdu/OpenOCR OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta… | 58 | 1437 | active |
| potatameister/PaperKnife PaperKnife is a privacy-first PDF utility that merges, splits, compresses, encrypts, signs, and sanitizes PDFs entirely on-device with no s… | 67 | 1435 | active |
| xu-cheng/latex-action A GitHub Action that compiles LaTeX documents to PDF inside a Dockerized full TeXLive environment. It supports multiple LaTeX engines, TeXL… | 86 | 1410 | active |
| mufeedvh/pdfrip pdfrip is a multithreaded PDF password cracking utility written in Rust. It supports dictionary attacks, mask and pattern-based brute force… | 68 | 1406 | active |
| alibaba/Logics-Parsing Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul… | 54 | 1402 | active |
| agentcooper/react-pdf-highlighter A set of React components for annotating PDF documents, built on top of PDF.js. It supports text and image highlights, popover text for hig… | 32 | 1401 | active |
| Skardyy/mcat Mcat is a Rust-based terminal tool that parses, converts, and previews files such as images, videos, PDFs, DOCX, HTML, and Markdown directl… | 85 | 1388 | active |
| 14790897/handwriting-web A self-hostable web application that converts typed text into images (or PDFs) simulating handwriting, using uploaded fonts, background ima… | 92 | 1385 | active |
| lamm-mit/PDF2Audio A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text ge… | 30 | 1383 | active |
| gettalong/hexapdf HexaPDF is a pure Ruby library with an accompanying CLI application for creating, manipulating, merging, encrypting, signing, and optimizin… | 76 | 1382 | active |
| SUSYUSTC/MathTranslate MathTranslate is a Python tool that translates LaTeX documents, especially scientific papers from arXiv, between any languages while keepin… | 41 | 1363 | active |
| Jaspersoft/jasperreports JasperReports is the world's most popular open-source Java reporting engine, capable of producing pixel-perfect documents from any kind of … | 94 | 1356 | active |
| mzucker/noteshrink A Python command-line script that cleans up scans and photos of handwritten notes by separating background from ink, quantizing colors, and… | 32 | 4841 | maintenance |
| VadimDez/ng2-pdf-viewer ng2-pdf-viewer is an Angular component library for rendering PDF files in web applications, built on top of PDF.js. It supports features li… | 64 | 1349 | active |
| MiniGlome/Archive.org-Downloader A Python 3 command-line script that downloads borrowable books from archive.org and Open Library and assembles them into PDF files. It requ… | 76 | 1348 | active |
| huridocs/pdf-document-layout-analysis A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl… | 82 | 1346 | active |
| CoderWanFeng/python-office python-office is a Python office-automation library that wraps PDF, Word, Excel, PPT, image, video, email, WeChat, and OCR operations behin… | 67 | 1344 | active |
| Swati4star/Images-to-PDF An open-source Android app that converts images (JPG and others) from the camera or gallery into PDF files. It also offers PDF management f… | 70 | 1328 | active |
| mpdf/mpdf mPDF is a PHP library that generates PDF documents from UTF-8 encoded HTML, based on FPDF and HTML2FPDF with extensive enhancements. It sup… | 62 | 4700 | maintenance |
| deanmalmgren/textract A Python library that extracts text from virtually any document format (PDF, DOCX, PPTX, HTML, images, and more) through a simple unified i… | 93 | 4698 | maintenance |
| gavrielc/Nano-PDF A Python CLI tool that edits PDF slides using natural language prompts, powered by Google's Gemini 3 Pro Image model. It renders pages to i… | 41 | 1321 | active |
| jsreport/jsreport jsreport is an open-source JavaScript-based reporting platform and server for designing and rendering reports using templating engines like… | 94 | 1319 | stable |
| yzane/vscode-markdown-pdf A Visual Studio Code extension that converts Markdown files to PDF, HTML, PNG, or JPEG. It supports PlantUML and Mermaid diagrams, KaTeX ma… | 93 | 1319 | active |
| shift-labs-ai/markit A Rust-based tool that converts documents, data files, web pages, and media into Markdown, usable both as a CLI and a library. It supports … | 81 | 1315 | active |
| CBIhalsen/PolyglotPDF PolyglotPDF is a Python-based multilingual eBook and PDF translation tool that preserves original layouts while translating, supporting bot… | 42 | 1315 | active |
| ndl-lab/ndlocr-lite NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma… | 77 | 1309 | active |
| SSShooter/ebook-to-mindmap A browser-based AI application that parses EPUB and PDF ebooks and converts them into chapter summaries, per-chapter mind maps, or a single… | 72 | 1306 | active |
| ttop32/MouseTooltipTranslator A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports … | 95 | 1300 | active |
| wisupai/e2m E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using… | 23 | 1294 | active |
| mb21/panwriter PanWriter is a free, open-source, distraction-free Markdown editor with tight pandoc integration for importing and exporting many document … | 65 | 1292 | active |
| fraserxu/electron-pdf Electron-PDF is a command-line tool and Node.js library that uses Electron's Chromium renderer to generate PDF files from URLs, HTML, or Ma… | 53 | 1292 | active |
| xunbu/docutranslate DocuTranslate is a lightweight local document translation tool powered by large language models, supporting formats such as pdf, docx, xlsx… | 80 | 1285 | active |
| jbarrow/commonforms CommonForms is a Python package and CLI that uses trained object-detection models (FFDNet-S/L) to automatically detect form fields in a PDF… | 65 | 1284 | active |
| Leseratte10/acsm-calibre-plugin A Calibre FileType plugin that converts ACSM files into EPUB or PDF without needing Adobe Digital Editions, implemented as a full Python re… | 65 | 1282 | active |
| dvcoolarun/web2pdf A Python command-line tool that converts webpages into formatted PDFs using WeasyPrint. It supports batch conversion, recursive same-domain… | 67 | 1281 | active |
| BHOSC/BUAAthesis BUAAthesis is a LaTeX thesis template for Beihang University (BUAA) graduation theses, maintained by the BUAA Open Source Club. It supports… | 23 | 1269 | active |
| Yuliang-Liu/MonkeyOCRv2 MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2… | 58 | 1256 | active |
| Setasign/FPDI FPDI is a collection of PHP classes that read pages from existing PDF documents and use them as templates in PDF generation libraries like … | 91 | 1248 | stable |
| KnpLabs/KnpSnappyBundle A Symfony bundle integrating the Snappy PHP wrapper around wkhtmltopdf/wkhtmltoimage, letting Symfony apps convert HTML documents or URLs i… | 70 | 1247 | active |
| gfngfn/SATySFi SATySFi is a typesetting system built around a statically-typed, functional programming language, producing PDF documents. It combines a La… | 57 | 1247 | active |
| DDULDDUCK/every-pdf Every PDF is an all-in-one open-source desktop PDF toolkit built with Electron, Next.js, and a Python (FastAPI) backend. It lets users edit… | 71 | 1243 | active |
| jlegewie/zotfile ZotFile is a Zotero plugin for managing PDF attachments: it automatically renames, moves, and attaches PDFs to Zotero items, syncs PDFs to … | 23 | 4367 | maintenance |
| chinapandaman/PyPDFForm PyPDFForm is a Python library and CLI tool for creating, inspecting, styling, and filling PDF forms, along with common PDF utilities like p… | 93 | 1242 | active |