Ross ROSS = Recommend OSS · open-source software intelligence for agents

MarkPDFdown/markpdfdown

A high-quality PDF to Markdown tool based on large language model visual recognition. 一款基于大模型视觉识别的高质量PDF转Markdown工具 observed · 2026-08-28

github.com/MarkPDFdown/markpdfdown · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

60/100

  • Activity 64
  • Release rhythm 67
  • Longevity 38
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 2.0
  • age_days: 537
  • days_rel: 220
  • days_push: 220
  • n_releases_24m: 13

Full methodology

Adoption not part of the score

2184 stars · 180 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

MarkPDFDown is a Python CLI tool that converts PDF documents and images into clean Markdown using multimodal large language models via LiteLLM. It preserves formatting such as headings, tables, formulas, and diagrams, and supports OpenAI and OpenRouter providers.

Use cases

  • convert pdf to markdown
  • extract tables and formulas from pdfs into markdown
  • transcribe scanned documents with a vision llm
  • convert images to markdown text
  • convert only specific page ranges of a pdf
  • batch convert pdfs from the command line

When to choose

  • you need high-fidelity PDF-to-Markdown conversion of complex layouts, tables, and formulas
  • you already have an OpenAI or OpenRouter API key and want LLM-based extraction
  • you want a scriptable CLI or pipe-based conversion workflow

When to avoid

  • you need fully offline conversion without sending documents to an LLM API
  • you want a free tool with no per-page API costs
  • you only need simple text extraction that a lightweight parser could handle

Facets

cli-tool · maturity active

ocr pdf llm-inference cli pdf developer-tools large-language-models python cli cross-platform pdf-to-markdown multimodal-llm document-conversion litellm vision-models docker

1 source

Member repositories

RepositoryRoleHealth v2
MarkPDFdown/markpdfdownmain60

For agents

markdown · JSON · MCP: product_card(name="MarkPDFdown/markpdfdown")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem