domain: pdf
514 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| apache/tika Apache Tika is a Java toolkit that detects file types and extracts text and metadata from over a thousand file formats (PDF, Office documen… | 77 | 4011 | stable |
| HKUDS/Paper2Slides Paper2Slides is a Python application that converts research papers, reports, and documents into professional presentation slides and poster… | 53 | 3822 | active |
| camelot-dev/camelot Camelot is a Python library for extracting tabular data from PDFs, offering five parsers including heuristic (lattice, stream), text-alignm… | 94 | 3811 | active |
| bgreenwell/doxx doxx is a fast, terminal-native viewer for .docx Word documents built in Rust with Ratatui. It renders formatting, tables, lists, images, a… | 71 | 3747 | active |
| SteveTheKiller/KillerPDF KillerPDF is a free, open-source (GPL-3.0) PDF editor for Windows built with C# and WPF, offering viewing, annotation, OCR, form filling, s… | 81 | 3656 | active |
| lookscanned/lookscanned.io LookScanned.io is a pure frontend web application that transforms PDFs (and other documents) so they look like scanned copies, with adjusta… | 38 | 3590 | active |
| mileszs/wicked_pdf Wicked PDF is a Ruby on Rails plugin that generates PDF files from HTML views by wrapping the wkhtmltopdf shell utility. Developers write n… | 38 | 3572 | stable |
| borb-pdf/borb borb is a pure Python library for reading, creating, and manipulating PDF files, modeling documents in a JSON-like structure with a high-le… | 94 | 3570 | active |
| CosmosShadow/gptpdf A small Python library that parses PDF files into Markdown using a vision-capable LLM such as GPT-4o. It uses PyMuPDF to detect non-text ar… | 31 | 3564 | active |
| wkhtmltopdf/wkhtmltopdf wkhtmltopdf and wkhtmltoimage are open-source (LGPLv3) command line tools that render HTML into PDF and various image formats using the Qt … | 10 | 14572 | maintenance |
| deepseek-ai/DeepSeek-OCR-2 DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn… | 44 | 3379 | active |
| pwmt/zathura Zathura is a highly customizable, keyboard-driven document viewer built on the girara UI library and GTK4, with a plugin-based system for s… | 77 | 3260 | active |
| deepdoctection/deepdoctection deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c… | 98 | 3248 | active |
| shiyi-0x7f/o-lib Olib is a free, open-source cross-platform desktop e-book client built with Tauri 2 (Rust backend + React/TypeScript frontend). It provides… | 76 | 3232 | active |
| breezedeus/Pix2Text Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma… | 99 | 3227 | active |
| CatchTheTornado/text-extract-api A self-hosted FastAPI-based API that converts PDFs, Office documents, and images into Markdown or structured JSON using OCR engines (EasyOC… | 45 | 3175 | active |
| Filimoa/open-parse Open Parse is a Python library that visually parses complex documents (primarily PDFs) into semantically meaningful chunks for LLM and RAG … | 64 | 3159 | active |
| wdcpclover/ai4paper AI4Paper is an AI research workbench delivered primarily as a Zotero 7-10 plugin, with a web app, browser extension, and WeChat mini-progra… | 78 | 3142 | active |
| unidoc/unipdf UniPDF is a pure Go library for creating, reading, and processing PDF files, including report generation, form filling, digital signatures,… | 92 | 3106 | active |
| apache/pdfbox Apache PDFBox is an open source Java library for working with PDF documents, supporting creation, manipulation, and content extraction. It … | 77 | 3105 | stable |
| FastReports/FastReport FastReport is a free, open-source, band-oriented report generator library for .NET 6/.NET Core/.NET Framework written in C#. It lets applic… | 83 | 3086 | active |
| AnyListen/tools-ocr Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model… | 23 | 3065 | active |
| Dicklesworthstone/llm_aided_ocr A Python tool that converts scanned PDFs to text via Tesseract OCR, then uses LLMs (local or API-based like OpenAI/Anthropic) to correct OC… | 71 | 2993 | active |
| zcaceres/markdownify-mcp A Model Context Protocol (MCP) server that converts PDFs, images, audio, Office documents, and web content (including YouTube transcripts a… | 81 | 2983 | active |
| NVIDIA/NeMo-Retriever NVIDIA's scalable document content and metadata extraction library (also known as NVIDIA Ingest) that splits documents, classifies and extr… | 81 | 2970 | active |
| microsoft/table-transformer Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un… | 23 | 2939 | active |
| ciur/papermerge Papermerge is an open source document management system (DMS) designed for scanned documents and digital archives. It performs OCR on PDF, … | 59 | 2934 | active |
| signintech/gopdf gopdf is a Go library for programmatically generating PDF documents. It supports Unicode font embedding, drawing shapes and images, text al… | 75 | 2932 | active |
| ArtifexSoftware/mupdf MuPDF is a lightweight open-source C library and toolkit for viewing, rendering, parsing, and converting PDF, XPS, and e-book documents. It… | 77 | 2930 | stable |
| kane50613/takumi Takumi is a Rust-based rendering engine that converts JSX, HTML, and CSS into images (PNG, JPEG, WebP, SVG, animated GIF/WebP) and paged PD… | 81 | 2886 | active |
| MuiseDestiny/zotero-reference A Zotero add-on that extracts and displays the references of the PDF you are reading in a floating panel. It supports multiple data sources… | 75 | 2799 | active |
| pikepdf/pikepdf pikepdf is a Python library for reading, writing, repairing, and transforming PDF files, built on the mature qpdf C++ library. It supports … | 95 | 2798 | active |
| imanoop7/Ollama-OCR A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports… | 26 | 2780 | active |
| yilewang/llm-for-zotero A Zotero plugin that integrates large language models into the Zotero PDF reader, enabling grounded paper chat, summarization, figure extra… | 78 | 2761 | active |
| echohive42/AI-reads-books-page-by-page A Python script that reads PDF books page by page, using AI to extract knowledge points and generate progressive summaries at configurable … | 61 | 2756 | active |
| barryvdh/laravel-snappy A Laravel service provider wrapping the KnpLabs Snappy library to generate PDFs and images from HTML via wkhtmltopdf/wkhtmltoimage binaries… | 79 | 2753 | active |
| johnfercher/maroto Maroto is a Go library for generating PDFs using a Bootstrap-inspired grid system of rows, columns, and components, built on top of gofpdf.… | 98 | 2748 | active |
| freedmand/semantra-python Semantra is a command-line tool that semantically indexes text and PDF documents and launches a local web application for querying them by … | 30 | 2711 | active |
| naiveHobo/InvoiceNet InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN… | 32 | 2694 | active |
| chrome-php/chrome A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa… | 89 | 2675 | active |
| Ontos-AI/knowhere Knowhere is a document ingestion and parsing platform that transforms unstructured documents (PDF, DOCX, XLSX, PPTX, images) into structure… | 81 | 2673 | active |
| run-llama/sec-insights SEC Insights is a full-stack reference application built by LlamaIndex that uses RAG to answer questions about SEC 10-K and 10-Q financial … | 33 | 2609 | active |
| react-pdf-viewer/react-pdf-viewer A React component library for viewing PDF documents in the browser, built with TypeScript and React hooks on top of PDF.js. It offers a plu… | 10 | 2608 | active |
| bookfere/Ebook-Translator-Calibre-Plugin A Calibre plugin that translates ebooks into a specified language using engines like Google Translate, ChatGPT, Gemini, and DeepL, optional… | 83 | 2599 | active |
| lncrawl/lightnovel-crawler Lightnovel Crawler is a Python tool that downloads web novels from 300+ supported sources and converts them into e-books such as EPUB, MOBI… | 97 | 2575 | active |
| sismics/docs Teedy (formerly Sismics Docs) is an open-source, lightweight document management system (DMS) for individuals and businesses. It offers OCR… | 67 | 2561 | active |
| OnedocLabs/react-print-pdf react-print-pdf is an open-source React component library for building and generating PDF documents using reusable, unstyled components and… | 26 | 2556 | active |
| UglyToad/PdfPig PdfPig is a C#/.NET library for reading and extracting text, images, annotations, forms, and metadata from PDF files, ported from Apache PD… | 91 | 2549 | active |
| simonbengtsson/jsPDF-AutoTable A jsPDF plugin that generates PDF tables in JavaScript, either from HTML table elements or from JavaScript data. It supports themes, column… | 81 | 2541 | stable |
| chatdoc-com/OCRFlux OCRFlux is a Python toolkit built on a 3B-parameter vision-language model that converts PDFs and images into clean Markdown, handling compl… | 53 | 2533 | active |
| facebookresearch/nougat Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py… | 23 | 10063 | maintenance |
| pipwerks/PDFObject PDFObject is a lightweight JavaScript utility for dynamically embedding PDF files into HTML documents using iframes. It detects browser PDF… | 31 | 2500 | stable |
| openpaperwork/paperwork Paperwork is a personal document manager for Linux and Windows that scans, OCRs, indexes, and organizes paper documents. It provides keywor… | 10 | 2432 | active |
| astefanutti/decktape Decktape is a command-line tool that converts HTML presentation frameworks (like reveal.js, impress.js, and others) into high-quality PDF d… | 78 | 2422 | active |
| RyotaUshio/obsidian-pdf-plus PDF++ is an Obsidian.md community plugin that upgrades Obsidian's built-in PDF viewer into a full annotation and viewing tool. It lets user… | 47 | 2421 | active |
| xhtml2pdf/xhtml2pdf xhtml2pdf is a pure-Python library that converts HTML and CSS (HTML5, CSS 2.1, partial CSS 3) into PDF documents using ReportLab, html5lib,… | 67 | 2387 | active |
| plutext/docx4j docx4j is an open-source (Apache v2) Java library for creating, editing, and saving Microsoft OpenXML packages, including Word docx, PowerP… | 77 | 2380 | active |
| hithesis/hithesis hithesis is a LaTeX thesis template package for Harbin Institute of Technology covering all three campuses. It supports bachelor, master, a… | 88 | 2379 | active |
| snapotter-hq/SnapOtter SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities… | 80 | 2362 | active |
| hydropix/TranslateBooksWithLLMs A desktop application that translates full-length books, subtitles, and documents using local or cloud LLMs (Ollama, OpenAI-compatible APIs… | 82 | 2334 | active |
| SwiftLaTeX/SwiftLaTeX SwiftLaTeX compiles the PdfTeX and XeTeX LaTeX engines to WebAssembly so LaTeX documents can be compiled entirely in the browser, exposed a… | 23 | 2315 | active |
| chezou/tabula-py tabula-py is a Python wrapper around tabula-java that extracts tables from PDF files into pandas DataFrames. It can also convert PDFs direc… | 23 | 2315 | stable |
| opendatalab/DocLayout-YOLO DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d… | 29 | 2258 | active |
| iText iText is a high-performance PDF library/SDK for Java and .NET that lets developers create, manipulate, inspect, sign, and secure PDF docume… | 93 | 2255 | stable |
| J-F-Liu/lopdf lopdf is a Rust library for low-level PDF document manipulation, supporting creating, parsing, editing, and merging PDF files. It offers op… | 97 | 2236 | stable |
| flyingsaucerproject/flyingsaucer Flying Saucer is a pure-Java library that renders well-formed XML/XHTML using CSS 2.1 layout and formatting. It outputs rendered documents … | 98 | 2235 | active |
| wxyhgk/retain-pdf RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/… | 77 | 2222 | active |
| Hopding/pdf-lib pdf-lib is a pure JavaScript/TypeScript library for creating and modifying PDF documents in any JavaScript runtime, including Node, browser… | 23 | 8605 | maintenance |
| modesty/pdf2json A Node.js module that converts binary PDF files into structured JSON and text, built on top of Mozilla's pdf.js. It extracts text content a… | 78 | 2210 | active |
| invoice-x/invoice2data A Python library and CLI tool that extracts structured data from PDF invoices using pluggable text-extraction backends (including OCR) and … | 97 | 2203 | active |
| MarkPDFdown/markpdfdown MarkPDFDown is a Python CLI tool that converts PDF documents and images into clean Markdown using multimodal large language models via Lite… | 60 | 2184 | active |
| TimmyOVO/deepseek-ocr.rs A Rust implementation of the DeepSeek-OCR inference stack with multiple OCR/VLM backends (DeepSeek-OCR, PaddleOCR-VL, DotsOCR), DSQ quantiz… | 60 | 2182 | active |
| ningzimu/image-to-editable-ppt-skill A Codex skill that converts slide images, PDFs, and image-based PPTX files into fully editable PowerPoint decks. It normalizes inputs into … | 76 | 2180 | active |
| scambier/obsidian-omnisearch Omnisearch is an Obsidian community plugin providing instant, relevance-weighted full-text search across notes, PDFs, Office documents, and… | 98 | 2127 | active |
| ustctug/ustcthesis A LaTeX template (ustcthesis) for writing undergraduate and graduate theses at the University of Science and Technology of China, conformin… | 92 | 2125 | active |
| GongRzhe/Office-Word-MCP-Server A Model Context Protocol (MCP) server that lets AI assistants create, read, format, and manipulate Microsoft Word (.docx) documents through… | 10 | 2107 | active |
| drunkdream/weread-exporter A Python CLI tool that exports books from WeChat Read (微信读书) into epub, pdf, and mobi formats. It hooks into the web reader's Canvas render… | 64 | 2101 | active |
| carboneio/carbone Carbone is a fast report and document generator that merges JSON data into templates created in Word, Excel, PowerPoint, LibreOffice, HTML,… | 66 | 2098 | active |
| NanoNets/docext docext is an on-premises document intelligence toolkit powered by vision-language models, offering OCR-free structured data extraction, PDF… | 50 | 2085 | active |
| raphaelmansuy/edgequake EdgeQuake is a high-performance Graph-RAG framework written in Rust, inspired by LightRAG, that transforms documents (PDFs, markdown, text)… | 79 | 2078 | active |
| WildDataX/suppr-zotero-plugin A Zotero plugin that translates academic PDF documents (and other formats via the Suppr web service) directly from the Zotero library via r… | 85 | 2011 | active |
| flytkgl/PDFQFZ PDFQFZ is a small desktop tool for adding cross-page (riding) seals to PDF documents. It takes a full seal image, randomly splits it across… | 90 | 2003 | active |
| manisandro/gImageReader gImageReader is a graphical GTK/Qt front-end to the tesseract-ocr engine for recognizing text in images, PDFs, scans, and screenshots. It s… | 51 | 1984 | active |
| Belval/pdf2image A Python module that wraps the pdftoppm and pdftocairo command-line utilities (from Poppler) to convert PDF pages into PIL Image objects. I… | 23 | 1982 | stable |
| run-llama/notebookllama NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval… | 52 | 1967 | active |
| SaiAkhil066/CORTEX-AI-SUPER-RAG CORTEX RAG is a local-first, agentic retrieval-augmented generation application that lets users upload PDFs and ask questions with cited an… | 61 | 1962 | active |
| Tabula Tabula is a local web application for extracting data tables from text-based PDF files into CSV, Excel, or JSON. It is powered by the tabul… | 28 | 7472 | maintenance |
| simonhaenisch/md-to-pdf A hackable Node.js CLI tool that converts Markdown files to PDF using Marked for HTML conversion and Puppeteer (headless Chromium) for rend… | 54 | 1950 | active |
| itsjunetime/tdf tdf is a terminal-based PDF viewer written in Rust using ratatui, designed for performance and responsiveness even with very large PDFs. It… | 76 | 1948 | active |
| clawsoftware/clawPDF clawPDF is an open-source virtual printer for Windows that converts printed output into PDF, PDF/A, OCR text, SVG, and various image format… | 23 | 1940 | active |
| Tencent-Hunyuan/HunyuanOCR HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model from Tencent, with a unified inference environment, llama.cpp PC-side … | 59 | 1930 | active |
| yob/pdf-reader PDF::Reader is a Ruby library that parses PDF files according to the Adobe PDF specification, providing programmatic access to document met… | 76 | 1926 | active |
| Lulzx/tinypdf tinypdf is a minimal TypeScript library for creating PDF documents with under 400 lines of code and zero dependencies. It supports text, sh… | 70 | 1908 | active |
| ranuts/document A browser-based document editor that opens and edits DOCX, XLSX, PPTX, CSV, PDF and other formats entirely client-side using the OnlyOffice… | 79 | 1902 | active |
| bhaskatripathi/pdfGPT pdfGPT is an open-source Python application that lets users chat with the contents of PDF files using GPT capabilities. It implements a sim… | 62 | 7167 | maintenance |
| rdumasia303/deepseek_ocr_app A self-hosted OCR web application combining a React frontend with a FastAPI backend, powered by the DeepSeek-OCR model. It processes images… | 50 | 1892 | active |
| alvarcarto/url-to-pdf-api A self-hosted microservice that converts URLs or raw HTML into PDF or PNG/JPEG images using Headless Chrome via Puppeteer. It exposes a sim… | 32 | 7105 | maintenance |
| pdfpc/pdfpc pdfpc is a GTK-based presenter console for PDF presentations with multi-monitor support. It shows the slides to the audience on one screen … | 58 | 1867 | active |
| run-llama/semtools SemTools is a Rust command-line toolkit for parsing documents (PDF, DOCX, PPTX, etc.) into markdown via LlamaParse and running fast local s… | 63 | 1858 | active |
| DeppWang/youdaonote-pull A Python script that exports and backs up all notes from Youdao Note (有道云笔记) to local storage. It converts note files to Markdown and downl… | 23 | 1855 | active |