function: pdf
749 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| TruthHun/DocHub DocHub is an open-source document library (wenku) system inspired by Baidu Wenku, built with Go and the Beego framework. It converts Office… | 23 | 2952 | maintenance |
| pdfkit/pdfkit PDFKit is a Ruby gem that converts HTML and CSS into PDF documents by wrapping the wkhtmltopdf command-line utility, which renders HTML via… | 23 | 2940 | maintenance |
| adieuadieu/serverless-chrome Serverless Chrome is a scaffolding framework and Serverless Framework plugin for running headless Chrome/Chromium on AWS Lambda. It bundles… | 23 | 2887 | maintenance |
| alanshaw/markdown-pdf A Node.js module and CLI that converts Markdown files to PDFs by rendering them through HTML and PhantomJS. It supports custom CSS styling,… | 32 | 2876 | maintenance |
| Xmader/musescore-downloader A tool for downloading sheet music from musescore.com for free without a login or Musescore Pro subscription, available as a CLI (npx msdl)… | 23 | 2762 | maintenance |
| smalot/pdfparser A standalone PHP library for parsing PDF files and extracting data such as text, metadata, and page content. It supports compressed PDFs an… | 82 | 2724 | maintenance |
| mangini/gdocs2md A Google Apps Script that converts a formatted Google Drive Document into Markdown, emailing the converted file with extracted images as at… | 32 | 2662 | maintenance |
| esbenp/pdf-bot A Node.js microservice that generates PDFs from URLs using headless Chrome, with a queue API, CLI, S3 storage, and webhook notifications. F… | 23 | 2641 | maintenance |
| cognitom/paper-css Paper CSS is a small CSS stylesheet that lets you create printable documents (A3/A4/A5, letter, legal) directly in the browser using plain … | 32 | 2586 | maintenance |
| JelteF/PyLaTeX PyLaTeX is a Python library for creating and compiling LaTeX files or snippets. It provides an easy but extensible interface between Python… | 23 | 2354 | maintenance |
| FranckFreiburger/vue-pdf A Vue.js 2 component that wraps Mozilla's PDF.js to render and display PDF documents in web applications. It exposes props for source, page… | 23 | 2298 | maintenance |
| danfickle/openhtmltopdf A pure-Java library that renders well-formed XHTML/HTML5 with CSS 2.1+ layout and outputs PDF documents (or images), built on Apache PDFBox… | 32 | 2176 | maintenance |
| jamiemcg/Remarkable Remarkable is a fully featured open-source Markdown editor for Linux with live preview, syntax highlighting, and GitHub Flavored Markdown s… | 23 | 2036 | maintenance |
| microsoft/maker.js Maker.js is a JavaScript/TypeScript library from Microsoft Garage for programmatically creating 2D vector line drawings using geometric and… | 67 | 2020 | maintenance |
| libharu/libharu libHaru is a free, open-source ANSI C library for generating PDF files, supporting text, lines, images, annotations, font embedding, compre… | 69 | 1993 | maintenance |
| webodf/ViewerJS ViewerJS is a JavaScript document viewer that renders PDF and OpenDocument Format (ODF) files directly in the browser, built on top of PDF.… | 23 | 1971 | maintenance |
| mdnice/markdown-resume An online resume formatting tool that lets users write resumes in Markdown or rich text and export them as a clean one-page PDF. It is a we… | 32 | 1970 | maintenance |
| alias-rahil/handwritten.js A JavaScript library, CLI tool, and web app that converts typed text into realistic-looking handwritten output rendered as PDF or image fil… | 32 | 1953 | maintenance |
| pmaupin/pdfrw pdfrw is a pure Python library that reads and writes PDF files, supporting operations like merging, rotating, subsetting, watermarking, and… | 23 | 1911 | maintenance |
| TCPDF TCPDF is a pure-PHP library for generating PDF documents and barcodes directly in application code, now in maintenance-only mode. It has be… | 77 | 1900 | maintenance |
| alejandro-ao/ask-multiple-pdfs A Streamlit-based Python application built with Langchain that lets users chat with multiple PDF documents using natural language. It extra… | 29 | 1873 | maintenance |
| ShizukuIchi/pdf-editor A browser-based PDF editor that lets users add images, signatures, and text to PDF files entirely client-side with no server or installatio… | 32 | 1864 | maintenance |
| kanishka-linux/reminiscence Reminiscence is a self-hosted bookmark and archive manager built with Django that saves web pages in HTML, PDF, or full-page PNG formats. I… | 23 | 1855 | maintenance |
| fossasia/badgeyay Badgeyay is a web application for generating attendee badges for conferences, concerts, and meetups. Users can customize badge size and bac… | 10 | 1794 | maintenance |
| there4/markdown-resume A command-line tool that converts Markdown resume documents into responsive HTML5/CSS3 resumes and printable PDFs via wkhtmltopdf. It offer… | 23 | 1785 | maintenance |
| impira/docquery DocQuery is a Python library and CLI tool that uses large language models to answer questions about semi-structured and unstructured docume… | 32 | 1775 | maintenance |
| mszep/pandoc_resume A template-driven resume builder that lets you write your resume once in Markdown and generate PDF, HTML, and other formats via pandoc and … | 32 | 1775 | maintenance |
| dbashford/textract A Node.js library that extracts plain text from many document formats including HTML, PDF, DOC/DOCX, XLS/XLSX, CSV, PPTX, RTF, EPUB, and im… | 58 | 1694 | maintenance |
| vladocar/screenshoteer Screenshoteer is a Node.js command-line tool built on Puppeteer that captures full-page website screenshots and emulates mobile devices. It… | 32 | 1667 | maintenance |
| jzillmann/pdf-to-markdown A JavaScript tool that parses PDF files and converts them into Markdown format, built on top of Mozilla's pdf.js. It is available as a host… | 64 | 1654 | maintenance |
| JonathanLink/PDFLayoutTextStripper A Java library that converts PDF files to text while preserving the original layout, built as a subclass of Apache PDFBox's PDFTextStripper… | 23 | 1608 | maintenance |
| politza/pdf-tools pdf-tools is an Emacs library that provides fast, feature-rich PDF viewing and interaction inside Emacs, backed by a C server (epdfinfo) fo… | 10 | 1574 | maintenance |
| BasioMeusPuga/Lector Lector is a Qt-based ebook reader application written in Python that supports multiple formats including PDF, EPUB, DjVu, MOBI, AZW, FB2, c… | 64 | 1551 | maintenance |
| benbalter/word-to-markdown A Ruby gem that converts Microsoft Word (.docx) documents into clean Markdown, using LibreOffice under the hood. It can be used as a librar… | 66 | 1550 | maintenance |
| liady/ChatGPT-pdf A Chrome and Firefox browser extension that adds export buttons to the ChatGPT web interface, letting users download chat history as PNG im… | 22 | 1480 | maintenance |
| ohoachuck/wwdc-downloader A Swift script for bulk downloading Apple WWDC session videos, PDF slides, and sample code in one shot, with no external dependencies. It s… | 71 | 1479 | maintenance |
| jonnnnyw/php-phantomjs A PHP library that wraps the PhantomJS headless browser, letting PHP applications load web pages with full JavaScript support and inspect t… | 23 | 1426 | maintenance |
| ArthurHub/HTML-Renderer A 100% managed C# library that renders HTML 4.01 and CSS level 2 content across WinForms, WPF, Mono, and PDF generation scenarios. It provi… | 97 | 1392 | maintenance |
| captn3m0/google-sre-ebook A shell-based generator that downloads the Google SRE books from sre.google and compiles them into EPUB, MOBI, and PDF ebook formats. Prebu… | 67 | 1391 | maintenance |
| nea/MarkdownViewerPlusPlus A Notepad++ plugin that renders Markdown files on-the-fly in a dockable preview panel. It supports synchronized scrolling, custom CSS, and … | 10 | 1386 | maintenance |
| nishiwen1214/ChatReviewer ChatReviewer is a Python application built on the ChatGPT API that analyzes academic papers, summarizing their strengths and weaknesses and… | 30 | 1377 | maintenance |
| prat0318/json_resume A Ruby CLI gem that converts a single JSON (or YAML) resume file into pretty HTML, PDF, LaTeX, and Markdown outputs using customizable temp… | 23 | 1364 | maintenance |
| SheetJS/js-word A pure-JavaScript parser and writer for word processing document formats including DOCX, DOC, and RTF, built from official specifications w… | 32 | 1321 | maintenance |
| empira/PDFsharp-1.5 PDFsharp is a .NET library for creating, modifying, and processing PDF documents programmatically, written in C#. It is often paired with M… | 32 | 1302 | maintenance |
| fathyb/html2svg A tool that renders web pages with a custom Chromium build and converts HTML and canvas content to SVG, PDF, or bitmap images (PNG, JPEG, W… | 23 | 1300 | maintenance |
| opensagres/xdocreport XDocReport is a Java API for merging XML documents created with MS Office (docx) or OpenOffice/LibreOffice (odt) with a Java model to gener… | 68 | 1299 | maintenance |
| JMCuixy/swagger2word A Java Spring Boot web application that converts Swagger/OpenAPI JSON specifications into downloadable Word documents. It accepts a Swagger… | 32 | 1281 | maintenance |
| mitchelloharawild/vitae An R package providing LaTeX and HTML templates plus helper functions for creating and maintaining résumés and CVs with R Markdown. It supp… | 87 | 1273 | maintenance |
| manuels/texlive.js A JavaScript port of TeX Live 2016 that compiles LaTeX code entirely in the browser using an emscripten-compiled pdftex. It produces PDF fi… | 32 | 1271 | maintenance |
| kayalshri/tableExport.jquery.plugin A jQuery plugin that exports HTML tables to multiple formats including JSON, XML, CSV, TXT, SQL, PNG, PDF, and Microsoft Office documents (… | 32 | 1235 | maintenance |
| mindbrix/UIImage-PDF An Objective-C UIImage category that renders PDF files into UIImages, enabling scalable vector assets in iOS apps. It supports rendering at… | 23 | 1226 | maintenance |
| rdvojmoc/DinkToPdf DinkToPdf is a C# .NET Core P/Invoke wrapper around the native wkhtmltopdf library, converting HTML pages or strings into PDF documents usi… | 23 | 1195 | maintenance |
| SebastiaanKlippert/go-wkhtmltopdf A pure Go wrapper around the wkhtmltopdf command-line utility for generating PDF documents from HTML/CSS templates. It exposes all wkhtmlto… | 45 | 1181 | maintenance |
| deepzec/Bad-Pdf Bad-PDF is a Python tool that generates malicious PDF files exploiting CVE-2018-4993 to steal NTLMv1/NTLMv2 hashes from Windows machines vi… | 44 | 1151 | maintenance |
| psiegman/epublib Epublib is a Java library for reading, writing, and manipulating EPUB ebook files, with a core that runs on both standard JVM and Android. … | 32 | 1123 | maintenance |
| josephernest/writing Writing is a lightweight, distraction-free Markdown and LaTeX text editor that runs entirely in the browser with no server code. It offers … | 32 | 1118 | maintenance |
| iainc/iA-Writer-Templates A collection of official and example templates for iA Writer that let users preview, export PDFs, and print documents in custom styles. Tem… | 32 | 1117 | maintenance |
| kevinhendricks/KindleUnpack A Python-based tool that unpacks non-DRM Amazon Kindle and MobiPocket ebooks into their component parts, such as HTML, images, PDFs, or epu… | 60 | 1097 | maintenance |
| metachris/pdfx PDFx is a Python command-line tool and library that extracts metadata, text, and references (PDFs, URLs, DOIs, arXiv IDs) from PDF document… | 10 | 1076 | maintenance |
| fuergaosi233/gitbook2pdf A Python CLI tool that asynchronously crawls GitBook documentation sites and converts their contents into a single PDF. It preserves the or… | 10 | 1067 | maintenance |
| jalan/pdftotext pdftotext is a Python library for simple PDF text extraction, built as a binding around the Poppler C++ library. It supports reading text p… | 10 | 1065 | maintenance |
| darwiin/yaac-another-awesome-cv YAAC: Another Awesome CV is a LaTeX class/template for building professional résumés and CVs, using XeLaTeX, Adobe Source Sans Pro, and Fon… | 23 | 1042 | maintenance |
| PDFHummus HummusJS is a fast Node.js native module (built on the PDF-Writer C++ library) for creating, parsing, and modifying PDF files and streams. … | 81 | 1021 | maintenance |
| IzakMarais/reporter An HTTP service written in Go that generates PDF reports from Grafana dashboards by rendering dashboard panels into a LaTeX-based PDF. It c… | 23 | 1021 | maintenance |
| DIYgod/Resume A simple HTML/CSS resume template that users fill in manually and deploy as a personal resume page. It supports generating a PDF via browse… | 32 | 1015 | maintenance |
| Algebra-FUN/WeReadScan A Python library that uses Selenium headless browsers to scan purchased books from WeRead (WeChat Reading) and convert them into local PDF … | 32 | 1002 | maintenance |
| dawnlabs/alchemy Alchemy is an open-source desktop file converter built with Electron and React that lives in the macOS/Windows menu bar. It lets users drag… | 23 | 1002 | maintenance |
| IlyaRice/RAG-Challenge-2 A Python implementation of the winning RAG system from the Enterprise RAG Challenge 2 competition, answering questions about company annual… | 29 | 2431 | experimental |
| cyphar/paperback paperback is a Rust CLI tool that creates encrypted paper backups suitable for long-term storage. It encrypts data and splits the secret ke… | 54 | 1478 | experimental |
| jlsutherland/doc2text doc2text is a Python library that extracts high-quality text from poorly scanned PDFs by correcting resolution, cropping, and skew before O… | 32 | 1278 | experimental |
| ThomasRinsma/pdftris A Python-generated PDF file that runs a playable Tetris game inside PDF viewers. It abuses PDF field objects to render a monochrome grid an… | 21 | 1038 | experimental |
| desgeeko/pdfsyntax PDFSyntax is a pure-Python, dependency-free library for inspecting and transforming the internal structure of PDF files down to the byte le… | 39 | 1009 | experimental |
| coolwanglu/pdf2htmlEX pdf2htmlEX is a command-line tool that converts PDF files into HTML while preserving text, fonts, and layout using modern web technologies.… | 10 | 10602 | abandoned |
| the-paperless-project/paperless A self-hosted application for scanning, OCR-indexing, and archiving paper documents with full-text search. This original repo is archived a… | 10 | 7912 | abandoned |
| yhatt/marp The classic Marp desktop app, a simple Markdown presentation writer for creating slide decks. It has been discontinued since 2017 and the r… | 10 | 7856 | abandoned |
| smuyyh/BookReader An Android e-book reader app ('任阅') for online novels and local txt/pdf/epub books, featuring 3D page-turning effects, bookshelf management… | 32 | 6936 | abandoned |
| axa-group/Parsr Parsr is a document parsing and extraction toolchain that transforms PDFs, images, docx, and eml files into clean, enriched structured data… | 56 | 6178 | abandoned |
| euske/pdfminer PDFMiner is a pure-Python library and CLI tool for parsing PDF documents and extracting text along with layout information such as fonts, p… | 10 | 5273 | abandoned |
| saurabhdaware/text-to-handwriting A web application that converts typed text into images that look like handwritten assignments, with customizable handwriting fonts, ink col… | 10 | 5040 | abandoned |
| jung-kurt/gofpdf A Go library that generates PDF documents with high-level support for text, drawing, and images. It is pure Go with no dependencies beyond … | 10 | 4469 | abandoned |
| run-llama/llama_cloud_services LlamaCloud Services provides cloud-hosted document parsing and knowledge agent tooling, including the LlamaParse document parser that conve… | 71 | 4262 | abandoned |
| karpathy/reader3 A lightweight, self-hosted EPUB reader web app that presents books one chapter at a time so users can easily copy text into an LLM and read… | 40 | 3842 | abandoned |
| SuxueCode/WechatBakTool A C#-based Windows GUI tool for backing up WeChat PC chat history by decrypting WeChat's local database and exporting messages. The project… | 56 | 3710 | abandoned |
| saadq/resumake.io Resumake is a free, open-source web application for automatically generating elegant LaTeX resumes without ads, accounts, or data collectio… | 72 | 3585 | abandoned |
| marcbachmann/node-html-pdf A Node.js library and CLI that converts HTML documents to PDF using the PhantomJS headless browser, supporting page sizes, headers, and foo… | 10 | 3578 | abandoned |
| mukulpatnaik/researchgpt A FastAPI-based research assistant that lets users have a conversation with any PDF research paper by extracting text, generating embedding… | 10 | 3531 | abandoned |
| ArtifexSoftware/pdf2docx A Python library and CLI for converting PDF files into editable DOCX (Word) documents, with layout, text, and table extraction built on PyM… | 84 | 3502 | abandoned |
| zhoubear/open-paperless Open Paperless is a self-hosted document management application for scanning, indexing, and archiving paper documents, built as a simplifie… | 32 | 2558 | abandoned |
| WZBSocialScienceCenter/pdftabextract pdftabextract is a Python library of tools for extracting tabular data from OCR-processed scanned PDFs ('sandwich PDFs' containing scanned … | 32 | 2255 | abandoned |
| arachnys/athenapdf Athena is a Docker-powered HTML-to-PDF conversion tool consisting of an Electron-based CLI and a Go microservice, designed as a drop-in rep… | 10 | 2245 | abandoned |
| JazzCore/python-pdfkit A Python 3 wrapper around the wkhtmltopdf utility that converts HTML, URLs, or strings into PDF documents using Webkit rendering. It expose… | 23 | 2045 | abandoned |
| RD17/ambar Ambar is an open-source, self-hosted document search engine with automated file system crawling, OCR, tagging, and instant Google-like full… | 10 | 1943 | abandoned |
| scantailor/scantailor Scan Tailor is an interactive post-processing tool for scanned pages, performing operations like page splitting, deskewing, border adjustme… | 10 | 1800 | abandoned |
| bughandler/cnki-downloader A small desktop tool for searching and downloading academic literature from CNKI (China National Knowledge Infrastructure). Its backend int… | 48 | 1762 | abandoned |
| SmartSchoolAI/ai-to-pptx Ai-to-pptx is an open-source web application (Vue frontend plus PHP backend) that uses large language models like DeepSeek to generate pres… | 10 | 1459 | abandoned |
| mobfarm/FastPdfKit FastPdfKit is an Objective-C static library for embedding fast PDF viewing in iOS applications, offering features like page sliding, search… | 32 | 1220 | abandoned |
| ZhongXiaoHong/superFileView An Android demo application that displays documents (doc, docx, xls, ppt, pdf, txt) using Tencent's TBS X5 WebKit kernel. It supports viewi… | 32 | 1198 | abandoned |
| xiaoguyu/wechatDownload A desktop application built with Electron, TypeScript, and Vue3 for downloading WeChat Official Account (公众号) articles. It captures require… | 10 | 1170 | abandoned |
| covidpass-org/covidpass CovidPass is a web application that converts EU Digital COVID Certificates (PDF or QR code screenshots) into wallet passes for Apple Wallet… | 23 | 1157 | abandoned |
| mikemaccana/python-docx A Python library for creating, reading, querying, and modifying Microsoft Word 2007/2008 .docx (Office Open XML) files, supporting paragrap… | 10 | 1075 | abandoned |