Ross ROSS = Recommend OSS · open-source software intelligence for agents

hanshuaikang/AI-Media2Doc

一键将音视频转化为小红书/公众号/知识笔记/思维导图/视频字幕等各种风格的文档。 observed · 2026-08-28

github.com/hanshuaikang/AI-Media2Doc · Vue · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

57/100

  • Activity 66
  • Release rhythm 58
  • Longevity 36
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 9.0
  • age_days: 508
  • days_rel: 283
  • days_push: 209
  • n_releases_24m: 19

Full methodology

Adoption not part of the score

3994 stars · 542 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A self-hostable web application that uses AI large language models to convert video and audio into various document styles such as Xiaohongshu posts, WeChat articles, knowledge notes, mind maps, and subtitles. It runs locally via Docker with no login required, using ffmpeg.wasm in the browser for media processing.

Use cases

  • convert video to text notes
  • generate subtitles from audio or video
  • turn a YouTube or Bilibili video into a blog article
  • create Xiaohongshu posts from video content
  • make a mind map from a lecture recording
  • summarize podcast audio into knowledge notes
  • ask AI questions about a video's content
  • self-host a media-to-document transcription tool

When to choose

  • you want to transcribe and restyle video/audio content into documents without paid accounts
  • privacy matters and you prefer local deployment with no third-party uploads
  • you want automatic smart screenshots inserted into generated articles
  • you need one-click subtitle export or customizable prompts

When to avoid

  • you need fully offline transcription today - it relies on cloud LLM APIs (local fast-whisper is only planned)
  • you need a managed SaaS with team collaboration features
  • you need batch processing of large media libraries at scale

Facets

application · maturity active

speech-recognition llm-inference video-processing audio-processing web-framework chat-interface prompt-engineering artificial-intelligence large-language-models media developer-tools self-hosted web-development self-hosted python cross-platform transcription subtitle-generation ffmpeg-wasm xiaohongshu note-taking mindmap bilibili youtube content-creation docker web-server

1 source

Member repositories

RepositoryRoleHealth v2
hanshuaikang/AI-Media2Docmain57

For agents

markdown · JSON · MCP: product_card(name="hanshuaikang/AI-Media2Doc")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem