adithya-s-k/omniparse
Ingest, parse, and optimize any data format ➡️ from documents to multimedia ➡️ for enhanced compatibility with GenAI frameworks observed · 2026-08-28
Health v2 · maintenance only
49/100
- Activity 56
- Release rhythm 35
- Longevity 58
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 820
- days_rel: n/a
- days_push: 264
- n_releases_24m: 0
Adoption not part of the score
7815 stars · 667 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into clean, structured markdown optimized for LLM applications. It runs as a local API server with a Gradio UI, supporting ~20 file types with OCR, table extraction, transcription, and web crawling.
Use cases
- parse resumes from pdfs
- convert documents to markdown for RAG
- transcribe audio and video files locally
- extract tables from images and documents
- crawl web pages into clean markdown for LLM training
- prepare unstructured data for fine-tuning
- self-hosted document ingestion pipeline
When to choose
- you need fully local, private data parsing with no external APIs
- you want one tool handling documents, multimedia, and web content
- you're building RAG or fine-tuning pipelines and need LLM-friendly structured output
- you can run on Linux with optional GPU acceleration
When to avoid
- you need Windows or macOS native support (server is Linux-only)
- you only need simple PDF text extraction without multimedia or web crawling
- you require a managed cloud service rather than self-hosting
- your project needs a permissive license (GPL-3.0)
Facets
service · maturity active
ocr parser web-scraping speech-recognition rag etl api-framework nlp large-language-models developer-tools pdf crawlers self-hosted python data-ingestion document-parsing multimedia-transcription genai-preprocessing markdown-conversion gradio-ui data-engineering natural-language-processing retrieval-augmented-generation linux docker web-server gpu
2 sources
- readme: https://github.com/adithya-s-k/omniparse · fetched 2026-08-28 · 9b75ce401b70
- homepage: https://omniparse.cognitivelab.in/ · fetched 2026-08-29 · 5960fc67039a
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| adithya-s-k/omniparse | main | 49 |
For agents
markdown · JSON · MCP: product_card(name="adithya-s-k/omniparse")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem