Ross ROSS = Recommend OSS · open-source software intelligence for agents

lamm-mit/PDF2Audio

None observed · 2026-08-28

github.com/lamm-mit/PDF2Audio · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

30/100

  • Activity 17
  • Release rhythm 35
  • Longevity 50

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 710
  • days_rel: n/a
  • days_push: 502
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1383 stars · 175 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A Gradio-based web application that converts PDF documents into audio podcasts, lectures, and summaries using OpenAI GPT models for text generation and text-to-speech. Users can upload multiple PDFs, choose instruction templates, select voices, and iteratively edit the generated transcript.

Use cases

  • convert a pdf into a podcast
  • turn research papers into audio lectures
  • generate audio summaries of documents
  • create a notebooklm-style podcast from pdfs
  • edit and refine a podcast transcript before generating audio
  • choose different voices for podcast speakers

When to choose

  • you want a ready-made web UI for turning PDFs into spoken audio
  • you already have an OpenAI API key and want GPT-based script generation
  • you want to iterate on transcripts with comments before rendering audio

When to avoid

  • you need fully offline or self-hosted TTS without API costs
  • you need a CLI or batch pipeline rather than an interactive app
  • you need non-OpenAI LLM or TTS providers

Facets

application · maturity active

tts nlp llm-inference pdf audio-processing pdf artificial-intelligence developer-tools python cross-platform self-hosted pdf-to-audio podcast-generation gradio openai text-to-speech lecture-generation hugging-face-spaces natural-language-processing audio web-server

1 source

Member repositories

RepositoryRoleHealth v2
lamm-mit/PDF2Audiomain30

For agents

markdown · JSON · MCP: product_card(name="lamm-mit/PDF2Audio")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem