Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: pdf

514 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
apache/tika
Apache Tika is a Java toolkit that detects file types and extracts text and metadata from over a thousand file formats (PDF, Office documen…
774011stable
HKUDS/Paper2Slides
Paper2Slides is a Python application that converts research papers, reports, and documents into professional presentation slides and poster…
533822active
camelot-dev/camelot
Camelot is a Python library for extracting tabular data from PDFs, offering five parsers including heuristic (lattice, stream), text-alignm…
943811active
bgreenwell/doxx
doxx is a fast, terminal-native viewer for .docx Word documents built in Rust with Ratatui. It renders formatting, tables, lists, images, a…
713747active
SteveTheKiller/KillerPDF
KillerPDF is a free, open-source (GPL-3.0) PDF editor for Windows built with C# and WPF, offering viewing, annotation, OCR, form filling, s…
813656active
lookscanned/lookscanned.io
LookScanned.io is a pure frontend web application that transforms PDFs (and other documents) so they look like scanned copies, with adjusta…
383590active
mileszs/wicked_pdf
Wicked PDF is a Ruby on Rails plugin that generates PDF files from HTML views by wrapping the wkhtmltopdf shell utility. Developers write n…
383572stable
borb-pdf/borb
borb is a pure Python library for reading, creating, and manipulating PDF files, modeling documents in a JSON-like structure with a high-le…
943570active
CosmosShadow/gptpdf
A small Python library that parses PDF files into Markdown using a vision-capable LLM such as GPT-4o. It uses PyMuPDF to detect non-text ar…
313564active
wkhtmltopdf/wkhtmltopdf
wkhtmltopdf and wkhtmltoimage are open-source (LGPLv3) command line tools that render HTML into PDF and various image formats using the Qt …
1014572maintenance
deepseek-ai/DeepSeek-OCR-2
DeepSeek-OCR 2 is an open-source vision-language model and inference toolkit implementing 'Visual Causal Flow' for optical character recogn…
443379active
pwmt/zathura
Zathura is a highly customizable, keyboard-driven document viewer built on the girara UI library and GTK4, with a plugin-based system for s…
773260active
deepdoctection/deepdoctection
deepdoctection is a Python library for Document AI that orchestrates document layout analysis, table recognition, OCR, and document/token c…
983248active
shiyi-0x7f/o-lib
Olib is a free, open-source cross-platform desktop e-book client built with Tauri 2 (Rust backend + React/TypeScript frontend). It provides…
763232active
breezedeus/Pix2Text
Pix2Text is an open-source Python tool that recognizes layouts, tables, math formulas (LaTeX), and text in images and converts them into Ma…
993227active
CatchTheTornado/text-extract-api
A self-hosted FastAPI-based API that converts PDFs, Office documents, and images into Markdown or structured JSON using OCR engines (EasyOC…
453175active
Filimoa/open-parse
Open Parse is a Python library that visually parses complex documents (primarily PDFs) into semantically meaningful chunks for LLM and RAG …
643159active
wdcpclover/ai4paper
AI4Paper is an AI research workbench delivered primarily as a Zotero 7-10 plugin, with a web app, browser extension, and WeChat mini-progra…
783142active
unidoc/unipdf
UniPDF is a pure Go library for creating, reading, and processing PDF files, including report generation, form filling, digital signatures,…
923106active
apache/pdfbox
Apache PDFBox is an open source Java library for working with PDF documents, supporting creation, manipulation, and content extraction. It …
773105stable
FastReports/FastReport
FastReport is a free, open-source, band-oriented report generator library for .NET 6/.NET Core/.NET Framework written in C#. It lets applic…
833086active
AnyListen/tools-ocr
Tree Hole OCR is a cross-platform desktop OCR tool built with Java and JavaFX that performs offline text recognition using Paddle OCR model…
233065active
Dicklesworthstone/llm_aided_ocr
A Python tool that converts scanned PDFs to text via Tesseract OCR, then uses LLMs (local or API-based like OpenAI/Anthropic) to correct OC…
712993active
zcaceres/markdownify-mcp
A Model Context Protocol (MCP) server that converts PDFs, images, audio, Office documents, and web content (including YouTube transcripts a…
812983active
NVIDIA/NeMo-Retriever
NVIDIA's scalable document content and metadata extraction library (also known as NVIDIA Ingest) that splits documents, classifies and extr…
812970active
microsoft/table-transformer
Table Transformer (TATR) is a deep learning object detection model from Microsoft for detecting, recognizing, and extracting tables from un…
232939active
ciur/papermerge
Papermerge is an open source document management system (DMS) designed for scanned documents and digital archives. It performs OCR on PDF, …
592934active
signintech/gopdf
gopdf is a Go library for programmatically generating PDF documents. It supports Unicode font embedding, drawing shapes and images, text al…
752932active
ArtifexSoftware/mupdf
MuPDF is a lightweight open-source C library and toolkit for viewing, rendering, parsing, and converting PDF, XPS, and e-book documents. It…
772930stable
kane50613/takumi
Takumi is a Rust-based rendering engine that converts JSX, HTML, and CSS into images (PNG, JPEG, WebP, SVG, animated GIF/WebP) and paged PD…
812886active
MuiseDestiny/zotero-reference
A Zotero add-on that extracts and displays the references of the PDF you are reading in a floating panel. It supports multiple data sources…
752799active
pikepdf/pikepdf
pikepdf is a Python library for reading, writing, repairing, and transforming PDF files, built on the mature qpdf C++ library. It supports …
952798active
imanoop7/Ollama-OCR
A Python package and Streamlit web app that performs OCR on images and PDFs using vision language models served through Ollama. It supports…
262780active
yilewang/llm-for-zotero
A Zotero plugin that integrates large language models into the Zotero PDF reader, enabling grounded paper chat, summarization, figure extra…
782761active
echohive42/AI-reads-books-page-by-page
A Python script that reads PDF books page by page, using AI to extract knowledge points and generate progressive summaries at configurable …
612756active
barryvdh/laravel-snappy
A Laravel service provider wrapping the KnpLabs Snappy library to generate PDFs and images from HTML via wkhtmltopdf/wkhtmltoimage binaries…
792753active
johnfercher/maroto
Maroto is a Go library for generating PDFs using a Bootstrap-inspired grid system of rows, columns, and components, built on top of gofpdf.…
982748active
freedmand/semantra-python
Semantra is a command-line tool that semantically indexes text and PDF documents and launches a local web application for querying them by …
302711active
naiveHobo/InvoiceNet
InvoiceNet is a deep neural network application with a GUI for extracting structured information from invoice documents in PDF, JPG, and PN…
322694active
chrome-php/chrome
A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa…
892675active
Ontos-AI/knowhere
Knowhere is a document ingestion and parsing platform that transforms unstructured documents (PDF, DOCX, XLSX, PPTX, images) into structure…
812673active
run-llama/sec-insights
SEC Insights is a full-stack reference application built by LlamaIndex that uses RAG to answer questions about SEC 10-K and 10-Q financial …
332609active
react-pdf-viewer/react-pdf-viewer
A React component library for viewing PDF documents in the browser, built with TypeScript and React hooks on top of PDF.js. It offers a plu…
102608active
bookfere/Ebook-Translator-Calibre-Plugin
A Calibre plugin that translates ebooks into a specified language using engines like Google Translate, ChatGPT, Gemini, and DeepL, optional…
832599active
lncrawl/lightnovel-crawler
Lightnovel Crawler is a Python tool that downloads web novels from 300+ supported sources and converts them into e-books such as EPUB, MOBI…
972575active
sismics/docs
Teedy (formerly Sismics Docs) is an open-source, lightweight document management system (DMS) for individuals and businesses. It offers OCR…
672561active
OnedocLabs/react-print-pdf
react-print-pdf is an open-source React component library for building and generating PDF documents using reusable, unstyled components and…
262556active
UglyToad/PdfPig
PdfPig is a C#/.NET library for reading and extracting text, images, annotations, forms, and metadata from PDF files, ported from Apache PD…
912549active
simonbengtsson/jsPDF-AutoTable
A jsPDF plugin that generates PDF tables in JavaScript, either from HTML table elements or from JavaScript data. It supports themes, column…
812541stable
chatdoc-com/OCRFlux
OCRFlux is a Python toolkit built on a 3B-parameter vision-language model that converts PDFs and images into clean Markdown, handling compl…
532533active
facebookresearch/nougat
Nougat is Meta's neural OCR model that parses academic PDFs into structured Markdown, understanding LaTeX math and tables. It ships as a Py…
2310063maintenance
pipwerks/PDFObject
PDFObject is a lightweight JavaScript utility for dynamically embedding PDF files into HTML documents using iframes. It detects browser PDF…
312500stable
openpaperwork/paperwork
Paperwork is a personal document manager for Linux and Windows that scans, OCRs, indexes, and organizes paper documents. It provides keywor…
102432active
astefanutti/decktape
Decktape is a command-line tool that converts HTML presentation frameworks (like reveal.js, impress.js, and others) into high-quality PDF d…
782422active
RyotaUshio/obsidian-pdf-plus
PDF++ is an Obsidian.md community plugin that upgrades Obsidian's built-in PDF viewer into a full annotation and viewing tool. It lets user…
472421active
xhtml2pdf/xhtml2pdf
xhtml2pdf is a pure-Python library that converts HTML and CSS (HTML5, CSS 2.1, partial CSS 3) into PDF documents using ReportLab, html5lib,…
672387active
plutext/docx4j
docx4j is an open-source (Apache v2) Java library for creating, editing, and saving Microsoft OpenXML packages, including Word docx, PowerP…
772380active
hithesis/hithesis
hithesis is a LaTeX thesis template package for Harbin Institute of Technology covering all three campuses. It supports bachelor, master, a…
882379active
snapotter-hq/SnapOtter
SnapOtter is an open-source, self-hosted file-processing suite offering 200+ tools across image, video, audio, PDF, and document modalities…
802362active
hydropix/TranslateBooksWithLLMs
A desktop application that translates full-length books, subtitles, and documents using local or cloud LLMs (Ollama, OpenAI-compatible APIs…
822334active
SwiftLaTeX/SwiftLaTeX
SwiftLaTeX compiles the PdfTeX and XeTeX LaTeX engines to WebAssembly so LaTeX documents can be compiled entirely in the browser, exposed a…
232315active
chezou/tabula-py
tabula-py is a Python wrapper around tabula-java that extracts tables from PDF files into pandas DataFrames. It can also convert PDFs direc…
232315stable
opendatalab/DocLayout-YOLO
DocLayout-YOLO is a real-time YOLO-v10-based model for detecting document layout elements (text blocks, tables, figures, etc.) in diverse d…
292258active
iText
iText is a high-performance PDF library/SDK for Java and .NET that lets developers create, manipulate, inspect, sign, and secure PDF docume…
932255stable
J-F-Liu/lopdf
lopdf is a Rust library for low-level PDF document manipulation, supporting creating, parsing, editing, and merging PDF files. It offers op…
972236stable
flyingsaucerproject/flyingsaucer
Flying Saucer is a pure-Java library that renders well-formed XML/XHTML using CSS 2.1 layout and formatting. It outputs rendered documents …
982235active
wxyhgk/retain-pdf
RetainPDF is an open-source PDF translation tool that preserves layout, formulas, and document structure, with special support for scanned/…
772222active
Hopding/pdf-lib
pdf-lib is a pure JavaScript/TypeScript library for creating and modifying PDF documents in any JavaScript runtime, including Node, browser…
238605maintenance
modesty/pdf2json
A Node.js module that converts binary PDF files into structured JSON and text, built on top of Mozilla's pdf.js. It extracts text content a…
782210active
invoice-x/invoice2data
A Python library and CLI tool that extracts structured data from PDF invoices using pluggable text-extraction backends (including OCR) and …
972203active
MarkPDFdown/markpdfdown
MarkPDFDown is a Python CLI tool that converts PDF documents and images into clean Markdown using multimodal large language models via Lite…
602184active
TimmyOVO/deepseek-ocr.rs
A Rust implementation of the DeepSeek-OCR inference stack with multiple OCR/VLM backends (DeepSeek-OCR, PaddleOCR-VL, DotsOCR), DSQ quantiz…
602182active
ningzimu/image-to-editable-ppt-skill
A Codex skill that converts slide images, PDFs, and image-based PPTX files into fully editable PowerPoint decks. It normalizes inputs into …
762180active
scambier/obsidian-omnisearch
Omnisearch is an Obsidian community plugin providing instant, relevance-weighted full-text search across notes, PDFs, Office documents, and…
982127active
ustctug/ustcthesis
A LaTeX template (ustcthesis) for writing undergraduate and graduate theses at the University of Science and Technology of China, conformin…
922125active
GongRzhe/Office-Word-MCP-Server
A Model Context Protocol (MCP) server that lets AI assistants create, read, format, and manipulate Microsoft Word (.docx) documents through…
102107active
drunkdream/weread-exporter
A Python CLI tool that exports books from WeChat Read (微信读书) into epub, pdf, and mobi formats. It hooks into the web reader's Canvas render…
642101active
carboneio/carbone
Carbone is a fast report and document generator that merges JSON data into templates created in Word, Excel, PowerPoint, LibreOffice, HTML,…
662098active
NanoNets/docext
docext is an on-premises document intelligence toolkit powered by vision-language models, offering OCR-free structured data extraction, PDF…
502085active
raphaelmansuy/edgequake
EdgeQuake is a high-performance Graph-RAG framework written in Rust, inspired by LightRAG, that transforms documents (PDFs, markdown, text)…
792078active
WildDataX/suppr-zotero-plugin
A Zotero plugin that translates academic PDF documents (and other formats via the Suppr web service) directly from the Zotero library via r…
852011active
flytkgl/PDFQFZ
PDFQFZ is a small desktop tool for adding cross-page (riding) seals to PDF documents. It takes a full seal image, randomly splits it across…
902003active
manisandro/gImageReader
gImageReader is a graphical GTK/Qt front-end to the tesseract-ocr engine for recognizing text in images, PDFs, scans, and screenshots. It s…
511984active
Belval/pdf2image
A Python module that wraps the pdftoppm and pdftocairo command-line utilities (from Poppler) to convert PDF pages into PIL Image objects. I…
231982stable
run-llama/notebookllama
NotebookLlaMa is an open-source, Python-based alternative to Google's NotebookLM, backed by LlamaCloud for document ingestion and retrieval…
521967active
SaiAkhil066/CORTEX-AI-SUPER-RAG
CORTEX RAG is a local-first, agentic retrieval-augmented generation application that lets users upload PDFs and ask questions with cited an…
611962active
Tabula
Tabula is a local web application for extracting data tables from text-based PDF files into CSV, Excel, or JSON. It is powered by the tabul…
287472maintenance
simonhaenisch/md-to-pdf
A hackable Node.js CLI tool that converts Markdown files to PDF using Marked for HTML conversion and Puppeteer (headless Chromium) for rend…
541950active
itsjunetime/tdf
tdf is a terminal-based PDF viewer written in Rust using ratatui, designed for performance and responsiveness even with very large PDFs. It…
761948active
clawsoftware/clawPDF
clawPDF is an open-source virtual printer for Windows that converts printed output into PDF, PDF/A, OCR text, SVG, and various image format…
231940active
Tencent-Hunyuan/HunyuanOCR
HunyuanOCR-1.5 is a lightweight end-to-end OCR vision-language model from Tencent, with a unified inference environment, llama.cpp PC-side …
591930active
yob/pdf-reader
PDF::Reader is a Ruby library that parses PDF files according to the Adobe PDF specification, providing programmatic access to document met…
761926active
Lulzx/tinypdf
tinypdf is a minimal TypeScript library for creating PDF documents with under 400 lines of code and zero dependencies. It supports text, sh…
701908active
ranuts/document
A browser-based document editor that opens and edits DOCX, XLSX, PPTX, CSV, PDF and other formats entirely client-side using the OnlyOffice…
791902active
bhaskatripathi/pdfGPT
pdfGPT is an open-source Python application that lets users chat with the contents of PDF files using GPT capabilities. It implements a sim…
627167maintenance
rdumasia303/deepseek_ocr_app
A self-hosted OCR web application combining a React frontend with a FastAPI backend, powered by the DeepSeek-OCR model. It processes images…
501892active
alvarcarto/url-to-pdf-api
A self-hosted microservice that converts URLs or raw HTML into PDF or PNG/JPEG images using Headless Chrome via Puppeteer. It exposes a sim…
327105maintenance
pdfpc/pdfpc
pdfpc is a GTK-based presenter console for PDF presentations with multi-monitor support. It shows the slides to the audience on one screen …
581867active
run-llama/semtools
SemTools is a Rust command-line toolkit for parsing documents (PDF, DOCX, PPTX, etc.) into markdown via LlamaParse and running fast local s…
631858active
DeppWang/youdaonote-pull
A Python script that exports and backs up all notes from Youdao Note (有道云笔记) to local storage. It converts note files to Markdown and downl…
231855active

← prev page 2 / 6 next →