Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: pdf

514 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
ofdrw/ofdrw
OFDRW is an open-source Java library for reading, writing, and manipulating OFD (Open Fixed-layout Document) files, a Chinese national stan…
931854active
mrmn2/PdfDing
PdfDing is a self-hosted PDF manager, viewer, and editor with a browser-based interface that syncs reading progress across devices. It supp…
101854active
clovaai/donut
Donut is an OCR-free end-to-end Transformer model for visual document understanding tasks such as document classification and information e…
236919maintenance
realdennis/md2pdf
An offline web application that converts Markdown files to PDF using a split-pane editor and browser print-to-PDF. It runs entirely client-…
611829stable
jacoblee93/fully-local-pdf-chatbot
A Next.js web application that lets users chat with uploaded PDF documents entirely locally, performing chunking, embedding, vector storage…
521817active
nazdridoy/kokoro-tts
A Python CLI text-to-speech tool built on the Kokoro-82M model that converts text, EPUB, PDF, and TXT inputs into natural-sounding speech w…
881816active
camelot-dev/excalibur
Excalibur is a self-hosted web interface for extracting tabular data from text-based PDFs, built on top of the Camelot PDF table extraction…
711815active
wonday/react-native-pdf
A React Native component that renders PDF documents on iOS and Android. It supports loading PDFs from URLs, blobs, local files, or assets w…
901809active
LibPDF-js/core
LibPDF is a modern TypeScript PDF library for parsing, modifying, generating, and signing PDF documents. It combines lenient parsing (like …
801803active
sajari/docconv
A Go library (with a companion docd CLI/HTTP service) that converts PDF, DOC, DOCX, XML, HTML, RTF, ODT, Pages documents and images into pl…
231788active
sile-typesetter/sile
SILE is a typesetting system that produces beautiful printed documents as PDF output, conceptually similar to TeX but written from scratch …
761783active
sainnhe/caj2pdf-qt
A cross-platform GUI application that converts CAJ, KDH, and NH files (CNKI academic document formats) to PDF, built with Qt on top of caj2…
231780active
neuml/paperai
paperai is an AI application for medical and scientific papers that runs bulk LLM inference and RAG pipelines over article repositories to …
671779active
chrisryugj/kordoc
kordoc is a TypeScript CLI and MCP server that parses Korean document formats (HWP 3.x/5.x, HWPX, HWPML, PDF, XLS/XLSX, DOCX, and images wi…
811776active
yobix-ai/extractous
Extractous is a fast Rust library for extracting content and metadata from unstructured documents like PDF, Word, Excel, HTML, CSV, and ema…
241772active
spipu/html2pdf
Html2Pdf is a PHP library that converts specially prepared HTML markup into PDF documents, built on top of TCPDF. It is designed for genera…
411769active
golbin/hop
HOP is an open-source desktop application for viewing and editing HWP/HWPX documents (the Korean Hangul word processor format) on macOS, Wi…
771766active
nguyenq/tess4j
Tess4J is a Java JNA wrapper for the Tesseract OCR API, enabling optical character recognition in Java applications. It supports TIFF, JPEG…
911757stable
community-archive/obsidian-zotero-integration
An Obsidian community plugin that inserts citations and bibliographies and imports notes and PDF annotations from Zotero into Obsidian. It …
541757active
Feather-2/Burner-X
Paper Burner X is a browser-based AI workstation for processing, translating, and analyzing academic documents like PDFs, DOCX, PPTX, and E…
511750active
LLM Sherpa
LLM Sherpa is a Python client library providing APIs for layout-aware PDF and document parsing to feed LLM applications, backed by the open…
181749active
jesselau76/ebook-GPT-translator
A Python CLI and GUI tool that translates ebooks (TXT, EPUB, DOCX, PDF, optional MOBI) using LLM providers like OpenAI, Azure, Codex, Claud…
631728active
zstmfhy/zlibrary-to-notebooklm
A Python CLI tool that automatically downloads books from Z-Library and uploads them to Google NotebookLM in one command. It uses Playwrigh…
431711active
baskerville/plato
Plato is a document reader application for Kobo e-ink e-readers, written in Rust. It supports PDF, EPUB, DJVU, CBZ, FB2, MOBI, XPS and TXT …
651694active
pdf-rs/pdf
A Rust library for reading, manipulating, and writing PDF files. Reading is fairly mature, while modifying and writing PDFs is still experi…
741691active
foliojs/fontkit
fontkit is an advanced font engine library for Node.js and the browser, used by PDFKit. It parses many font formats and provides glyph shap…
231668stable
chrismattmann/tika-python
Tika-Python is a Python client library for Apache Tika's REST services, enabling document parsing, text extraction, and MIME detection nati…
851666active
Cimbali/pympress
Pympress is a dual-screen PDF presentation tool that shows slides on a projector while giving the presenter a private view with notes, curr…
561647active
lmn1919/dompdf.js
dompdf.js is a pure-frontend DOM-to-PDF engine built with Rust, WebAssembly, and TypeScript that renders vector PDFs directly in the browse…
791620active
mikehaertl/phpwkhtmltopdf
A slim PHP wrapper around the wkhtmltopdf (and optionally wkhtmltoimage) command-line tools, exposing a clean OOP interface for creating PD…
531608stable
clusterzx/paperless-ai
A self-hosted web application that automatically analyzes, tags, and classifies documents in Paperless-ngx using LLMs via OpenAI-compatible…
735910maintenance
LYiHub/mad-professor-public
A desktop AI companion application for reading academic papers, built with PyQt6. It combines PDF parsing and translation, RAG-based retrie…
271604active
enoch3712/ExtractThinker
ExtractThinker is a Python document intelligence library that uses LLMs to extract and classify structured data from documents like PDFs, i…
451595active
jodconverter/jodconverter
JODConverter is a Java library that automates document conversions between office formats (e.g., docx, odt, xlsx, pdf) by driving LibreOffi…
481594stable
kotaro-kinoshita/yomitoku
YomiToku is an AI-powered document image analysis engine specialized for Japanese, providing full-text OCR, layout analysis, table structur…
871579active
elementdavv/internet_archive_downloader
A Chrome and Firefox browser extension that downloads borrowable books from Internet Archive (archive.org) and full-view books from HathiTr…
751565active
drmingler/docling-api
A self-hostable FastAPI backend service that converts documents (PDF, DOCX, PPTX, HTML, images, CSV, AsciiDoc, Markdown) into Markdown usin…
651557active
LaravelDaily/laravel-invoices
A Laravel PHP package for generating PDF invoice files from customizable parameters like items, taxes, discounts, and shipping. Invoices ca…
681553active
alpha-liu-01/SpeedyNote
SpeedyNote is a free, open-source, cross-platform note-taking application built in C++ with Qt, designed for stylus users who need fast PDF…
871544active
4lex4/scantailor-advanced
ScanTailor Advanced is an interactive post-processing tool for scanned pages, merging features from ScanTailor Featured and Enhanced while …
231532stable
NanoNets/docstrange
DocStrange is a Python library and tool that converts documents (PDF, DOCX, PPTX, XLSX, images, URLs) into Markdown, JSON, CSV, or HTML usi…
401531active
DavBfr/dart_pdf
A Dart/Flutter library for creating PDF documents, offering both a low-level PDF generation API and a Flutter-like widget system for high-l…
741527stable
kudrykv/latex-yearly-planner
A Go/LaTeX tool that generates yearly PDF planners designed for e-ink tablets like Supernote and ReMarkable. It produces pre-generated plan…
591525active
emcf/thepipe
thepipe is a Python package and API that extracts clean markdown, multimodal media, and structured data from tricky documents like PDFs, Wo…
581524active
oschwartz10612/poppler-windows
A repository that packages prebuilt Poppler PDF binaries (from conda-forge's poppler-feedstock) with their dependencies into convenient Win…
791522active
DocumindHQ/documind
Documind is an open-source Node.js library that uses AI/LLMs to extract structured JSON data from PDFs and other documents based on customi…
331521active
yamlresume/yamlresume
YAMLResume is a resume-as-code tool that lets you write your CV in YAML and compile it to beautifully typeset PDFs via LaTeX. It ships as a…
861504active
tjmlabs/ColiVara
ColiVara is a hosted (and self-hostable) retrieval API that stores, searches, and retrieves documents using visual embeddings generated by …
651486active
extend-hq/ui
Extend UI is an open-source React component library (distributed via the shadcn registry) for building document-heavy product interfaces: P…
581486active
KDE/okular
Okular is KDE's universal document viewer for formats including PDF, PostScript, EPUB, Comic Book, and images, with support for annotations…
771485stable
pagedjs/pagedjs
Paged.js is an open-source JavaScript library that paginates HTML content in the browser, polyfilling the W3C CSS Paged Media and Generated…
671485active
JakubMelka/PDF4QT
PDF4QT is an open-source PDF editor, viewer, and rendering library written in C++ with Qt, implementing PDF 2.0 functionality. It ships a C…
871459active
spatie/pdf-to-image
A PHP library that converts PDF files to images (jpg, png, webp) using Imagick and Ghostscript. It supports rendering single or multiple pa…
881457active
bblanchon/pdfium-binaries
A project that automatically builds and distributes pre-compiled binaries of Google's PDFium PDF rendering library for many platforms and C…
951456active
Open-Source-Legal/OpenContracts
OpenContracts is a self-hosted, MIT-licensed document intelligence platform that turns document repositories into a programmable citation g…
941451active
dadoonet/fscrawler
FSCrawler is a Java-based file system crawler that indexes binary documents (PDF, MS Office, Open Office) into Elasticsearch, tracking new,…
881450active
Nutlope/pdftochat
PDFToChat is an open-source Next.js web application that lets users upload PDFs and chat with them using AI. It uses Together AI's Mixtral …
681441active
Topdu/OpenOCR
OpenOCR is an open-source Python toolkit for general OCR research and applications, covering text detection and recognition, formula and ta…
581437active
potatameister/PaperKnife
PaperKnife is a privacy-first PDF utility that merges, splits, compresses, encrypts, signs, and sanitizes PDFs entirely on-device with no s…
671435active
xu-cheng/latex-action
A GitHub Action that compiles LaTeX documents to PDF inside a Dockerized full TeXLive environment. It supports multiple LaTeX engines, TeXL…
861410active
mufeedvh/pdfrip
pdfrip is a multithreaded PDF password cracking utility written in Rust. It supports dictionary attacks, mask and pattern-based brute force…
681406active
alibaba/Logics-Parsing
Logics-Parsing is an end-to-end document parsing model from Alibaba that converts document images into structured output using a single mul…
541402active
agentcooper/react-pdf-highlighter
A set of React components for annotating PDF documents, built on top of PDF.js. It supports text and image highlights, popover text for hig…
321401active
Skardyy/mcat
Mcat is a Rust-based terminal tool that parses, converts, and previews files such as images, videos, PDFs, DOCX, HTML, and Markdown directl…
851388active
14790897/handwriting-web
A self-hostable web application that converts typed text into images (or PDFs) simulating handwriting, using uploaded fonts, background ima…
921385active
lamm-mit/PDF2Audio
A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text ge…
301383active
gettalong/hexapdf
HexaPDF is a pure Ruby library with an accompanying CLI application for creating, manipulating, merging, encrypting, signing, and optimizin…
761382active
SUSYUSTC/MathTranslate
MathTranslate is a Python tool that translates LaTeX documents, especially scientific papers from arXiv, between any languages while keepin…
411363active
Jaspersoft/jasperreports
JasperReports is the world's most popular open-source Java reporting engine, capable of producing pixel-perfect documents from any kind of …
941356active
mzucker/noteshrink
A Python command-line script that cleans up scans and photos of handwritten notes by separating background from ink, quantizing colors, and…
324841maintenance
VadimDez/ng2-pdf-viewer
ng2-pdf-viewer is an Angular component library for rendering PDF files in web applications, built on top of PDF.js. It supports features li…
641349active
MiniGlome/Archive.org-Downloader
A Python 3 command-line script that downloads borrowable books from archive.org and Open Library and assembles them into PDF files. It requ…
761348active
huridocs/pdf-document-layout-analysis
A Dockerized microservice by HURIDOCS that performs PDF document layout analysis, OCR, and element segmentation/classification (texts, titl…
821346active
CoderWanFeng/python-office
python-office is a Python office-automation library that wraps PDF, Word, Excel, PPT, image, video, email, WeChat, and OCR operations behin…
671344active
Swati4star/Images-to-PDF
An open-source Android app that converts images (JPG and others) from the camera or gallery into PDF files. It also offers PDF management f…
701328active
mpdf/mpdf
mPDF is a PHP library that generates PDF documents from UTF-8 encoded HTML, based on FPDF and HTML2FPDF with extensive enhancements. It sup…
624700maintenance
deanmalmgren/textract
A Python library that extracts text from virtually any document format (PDF, DOCX, PPTX, HTML, images, and more) through a simple unified i…
934698maintenance
gavrielc/Nano-PDF
A Python CLI tool that edits PDF slides using natural language prompts, powered by Google's Gemini 3 Pro Image model. It renders pages to i…
411321active
jsreport/jsreport
jsreport is an open-source JavaScript-based reporting platform and server for designing and rendering reports using templating engines like…
941319stable
yzane/vscode-markdown-pdf
A Visual Studio Code extension that converts Markdown files to PDF, HTML, PNG, or JPEG. It supports PlantUML and Mermaid diagrams, KaTeX ma…
931319active
shift-labs-ai/markit
A Rust-based tool that converts documents, data files, web pages, and media into Markdown, usable both as a CLI and a library. It supports …
811315active
CBIhalsen/PolyglotPDF
PolyglotPDF is a Python-based multilingual eBook and PDF translation tool that preserves original layouts while translating, supporting bot…
421315active
ndl-lab/ndlocr-lite
NDLOCR-Lite is a lightweight Japanese OCR application developed by the National Diet Library that converts digitized images of books and ma…
771309active
SSShooter/ebook-to-mindmap
A browser-based AI application that parses EPUB and PDF ebooks and converts them into chapter summaries, per-chapter mind maps, or a single…
721306active
ttop32/MouseTooltipTranslator
A browser extension (Chrome, Edge, Firefox) that translates any text you hover over or select, showing an inline tooltip. It also supports …
951300active
wisupai/e2m
E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using…
231294active
mb21/panwriter
PanWriter is a free, open-source, distraction-free Markdown editor with tight pandoc integration for importing and exporting many document …
651292active
fraserxu/electron-pdf
Electron-PDF is a command-line tool and Node.js library that uses Electron's Chromium renderer to generate PDF files from URLs, HTML, or Ma…
531292active
xunbu/docutranslate
DocuTranslate is a lightweight local document translation tool powered by large language models, supporting formats such as pdf, docx, xlsx…
801285active
jbarrow/commonforms
CommonForms is a Python package and CLI that uses trained object-detection models (FFDNet-S/L) to automatically detect form fields in a PDF…
651284active
Leseratte10/acsm-calibre-plugin
A Calibre FileType plugin that converts ACSM files into EPUB or PDF without needing Adobe Digital Editions, implemented as a full Python re…
651282active
dvcoolarun/web2pdf
A Python command-line tool that converts webpages into formatted PDFs using WeasyPrint. It supports batch conversion, recursive same-domain…
671281active
BHOSC/BUAAthesis
BUAAthesis is a LaTeX thesis template for Beihang University (BUAA) graduation theses, maintained by the BUAA Open Source Club. It supports…
231269active
Yuliang-Liu/MonkeyOCRv2
MonkeyOCRv2 is a document-native vision encoder and visual-text foundation model for Document AI, pretrained on the 113M-image MonkeyDoc v2…
581256active
Setasign/FPDI
FPDI is a collection of PHP classes that read pages from existing PDF documents and use them as templates in PDF generation libraries like …
911248stable
KnpLabs/KnpSnappyBundle
A Symfony bundle integrating the Snappy PHP wrapper around wkhtmltopdf/wkhtmltoimage, letting Symfony apps convert HTML documents or URLs i…
701247active
gfngfn/SATySFi
SATySFi is a typesetting system built around a statically-typed, functional programming language, producing PDF documents. It combines a La…
571247active
DDULDDUCK/every-pdf
Every PDF is an all-in-one open-source desktop PDF toolkit built with Electron, Next.js, and a Python (FastAPI) backend. It lets users edit…
711243active
jlegewie/zotfile
ZotFile is a Zotero plugin for managing PDF attachments: it automatically renames, moves, and attaches PDFs to Zotero items, syncs PDFs to …
234367maintenance
chinapandaman/PyPDFForm
PyPDFForm is a Python library and CLI tool for creating, inspecting, styling, and filling PDF forms, along with common PDF utilities like p…
931242active

← prev page 3 / 6 next →