Ross ROSS = Recommend OSS · open-source software intelligence for agents

pdf2htmlEX/pdf2htmlEX

Convert PDF to HTML without losing text or format. observed · 2026-08-28

github.com/pdf2htmlEX/pdf2htmlEX · homepage · HTML · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

37/100

  • Activity 32
  • Release rhythm 8
  • Longevity 100

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 3145
  • days_rel: n/a
  • days_push: 412
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

5589 stars · 517 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

pdf2htmlEX is a command-line tool that converts PDF files into HTML while preserving text, fonts, and formatting using modern web technologies. It supports precise rendering, text selection, font embedding, and offers ~50 options for use cases like PDF preview and online publishing.

Use cases

  • convert pdf to html without losing text or format
  • publish pdf documents online as web pages
  • make pdf text selectable and searchable in the browser
  • embed pdf documents in a website without plugins
  • convert scanned or complex pdfs with fonts and math formulas to html
  • render pdf magazines or books for reading while downloading

When to choose

  • you need accurate, high-fidelity pdf-to-html conversion with preserved typography
  • you want text selection and font embedding in the output html
  • you need a scriptable cli tool with many output options
  • you want to publish pdfs as plugin-free interactive web documents

When to avoid

  • you need to extract plain text or structured data rather than visual html
  • your pdfs rely on Type 3 fonts, which are not yet supported
  • you need a maintained library API rather than a standalone converter
  • output html file size is a critical constraint without post-processing

Facets

cli-tool · maturity active

pdf parser ocr pdf web-development files windows cli pdf-to-html document-conversion font-embedding web-publishing text-preservation linux macos docker

2 sources

Member repositories

RepositoryRoleHealth v2
pdf2htmlEX/pdf2htmlEXmain37

For agents

markdown · JSON · MCP: product_card(name="pdf2htmlEX/pdf2htmlEX")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem