Ross ROSS = Recommend OSS · open-source software intelligence for agents

CBIhalsen/PolyglotPDF

(eBook,PDFs Translation) A multilingual eBook processing tool supporting all eBook formats. Features online and offline translation while preserving original layouts. Compatible with both scanned and digital PDFs. Elegant user interface. The world's highest-performing open-source layout-preserving eBook translator. observed · 2026-08-28

github.com/CBIhalsen/PolyglotPDF · homepage · Python · GPL-3.0 (copyleft) observed · 2026-08-28

Health v2 · maintenance only

42/100

  • Activity 44
  • Release rhythm 28
  • Longevity 63
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 38.0
  • age_days: 886
  • days_rel: 465
  • days_push: 339
  • n_releases_24m: 3

Full methodology

Adoption not part of the score

1315 stars · 198 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

PolyglotPDF is a Python-based multilingual eBook and PDF translation tool that preserves original layouts while translating, supporting both digital and scanned PDFs via OCR. It offers a web UI, online and offline translation via LLM APIs (DeepSeek, GPT-4o-mini, Qwen, Doubao) and traditional services (DeepL, Bing, YouDao), and handles math formulas and tables.

Use cases

  • translate a pdf document to another language while keeping the original layout
  • translate scanned pdfs with ocr
  • translate ebooks in any format preserving formatting
  • translate academic papers with math formulas and latex
  • batch translate large pdf reports with tables
  • self-host a pdf translation web service
  • translate pdfs offline without cloud services

When to choose

  • you need layout-preserving pdf translation rather than plain text extraction
  • you work with scanned pdfs or complex layouts like tables and formulas
  • you want a self-hosted tool with a choice of translation backends including LLMs
  • you need to translate ebooks in multiple formats beyond pdf

When to avoid

  • you only need raw text extraction from pdfs without translation
  • you need pixel-perfect handling of complex vector math formulas inside tables
  • you require a fully mature, bug-free tool for heavily styled or bold/colored mixed text
  • you want a lightweight cli-only one-liner translation utility

Facets

application · maturity active

pdf ocr nlp machine-learning llm-inference gui web-framework pdf files artificial-intelligence python cross-platform self-hosted pdf-translation ebook-translator layout-preserving pymupdf scanned-pdf latex-formulas deepseek openai ebook-formats natural-language-processing localization web-server docker

2 sources

Member repositories

RepositoryRoleHealth v2
CBIhalsen/PolyglotPDFmain42

For agents

markdown · JSON · MCP: product_card(name="CBIhalsen/PolyglotPDF")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem