# hanshuaikang/AI-Media2Doc

一键将音视频转化为小红书/公众号/知识笔记/思维导图/视频字幕等各种风格的文档。

Repository: https://github.com/hanshuaikang/AI-Media2Doc
Canonical: https://ross.abutalabs.com/products/ai-media2doc
Language: Vue
License: MIT
License Family: permissive
Topics: ai, python, vue, bilibili, chatgpt, openai, youtube, subtitles-generator, xiaohongshu
Last push: 2026-02-05T15:34:45+00:00

## Health v2 (maintenance only)
Score: 57/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 66, release rhythm 58, longevity 36
- inputs: {"age_days": 508, "days_push": 209, "days_rel": 283, "gap_med": 9.0, "n_releases_24m": 19}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3994, forks 542 (observed 2026-08-28T04:08:31.865296+00:00)

## What it is
A self-hostable web application that uses AI large language models to convert video and audio into various document styles such as Xiaohongshu posts, WeChat articles, knowledge notes, mind maps, and subtitles. It runs locally via Docker with no login required, using ffmpeg.wasm in the browser for media processing.

## Use cases
- convert video to text notes
- generate subtitles from audio or video
- turn a YouTube or Bilibili video into a blog article
- create Xiaohongshu posts from video content
- make a mind map from a lecture recording
- summarize podcast audio into knowledge notes
- ask AI questions about a video's content
- self-host a media-to-document transcription tool

## When to choose
- you want to transcribe and restyle video/audio content into documents without paid accounts
- privacy matters and you prefer local deployment with no third-party uploads
- you want automatic smart screenshots inserted into generated articles
- you need one-click subtitle export or customizable prompts

## When to avoid
- you need fully offline transcription today - it relies on cloud LLM APIs (local fast-whisper is only planned)
- you need a managed SaaS with team collaboration features
- you need batch processing of large media libraries at scale

## Facets
- artifact type: application
- maturity: active
- function: speech-recognition, llm-inference, video-processing, audio-processing, web-framework, chat-interface, prompt-engineering
- domain: artificial-intelligence, large-language-models, media, developer-tools, self-hosted, web-development
- platform: self-hosted, python, cross-platform
- tags: transcription, subtitle-generation, ffmpeg-wasm, xiaohongshu, note-taking, mindmap, bilibili, youtube, content-creation, docker, web-server

## Member repositories
- hanshuaikang/AI-Media2Doc (main) score 57

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:31.865296+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:24:12.132111+00:00, confidence not recorded.
  - readme: https://github.com/hanshuaikang/AI-Media2Doc (fetched 2026-08-28T04:08:31.865296+00:00, sha 552e7683404a)
- Data as of 2026-08-30T08:39:29.467469+00:00.
