Ross ROSS = Recommend OSS · open-source software intelligence for agents

wisupai/e2m

E2M converts various file types (doc, docx, epub, html, htm, url, pdf, ppt, pptx, mp3, m4a) into Markdown. It’s easy to install, with dedicated parsers and converters, supporting custom configs. E2M offers an all-in-one, flexible, and open-source solution. observed · 2026-08-28

github.com/wisupai/e2m · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

23/100

  • Activity 0
  • Release rhythm 35
  • Longevity 54

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 759
  • days_rel: n/a
  • days_push: 724
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1294 stars · 74 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

E2M is a Python library that parses and converts many file types (doc, docx, epub, html, url, pdf, ppt, pptx, mp3, m4a) into Markdown using a parser-converter architecture. It aims to produce high-quality text data for RAG pipelines and model training or fine-tuning.

Use cases

  • convert pdf to markdown
  • convert docx and pptx files to markdown for llm training
  • prepare clean text data for rag pipelines
  • extract text from epub and html into markdown
  • transcribe mp3 audio files to markdown text
  • batch convert documents for fine-tuning datasets

When to choose

  • you need a single Python library that handles many document formats as input
  • you want configurable parsers and converters for markdown output
  • you are building RAG or fine-tuning data pipelines from mixed file sources

When to avoid

  • you only need one specific format and a lighter single-purpose converter
  • you need a polished GUI application rather than a library
  • you require battle-tested enterprise-grade document conversion at scale

Facets

library · maturity active

parser pdf ocr speech-recognition llm-inference rag pdf files large-language-models developer-tools python cli cross-platform markdown-conversion document-parsing file-converter data-preparation fine-tuning-data natural-language-processing retrieval-augmented-generation

1 source

Member repositories

RepositoryRoleHealth v2
wisupai/e2mmain23

For agents

markdown · JSON · MCP: product_card(name="wisupai/e2m")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem