Zipstack/unstract
LLM-Driven Extraction of Unstructured Data — Built for API Deployments & ETL Pipeline Workflows observed · 2026-08-28
Health v2 · maintenance only
88/100
- Activity 99
- Release rhythm 86
- Longevity 66
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 0.0
- age_days: 924
- days_rel: 13
- days_push: 7
- n_releases_24m: 415
Adoption not part of the score
7172 stars · 709 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Unstract is an open-source, LLM-driven platform that extracts structured JSON data from unstructured documents such as PDFs, images, and scans using natural-language prompts. It supports deployment as APIs, ETL pipelines, and an MCP server, with features like dual-LLM validation (LLMChallenge) and token-saving extraction modes.
Use cases
- extract structured data from pdfs with llm
- parse invoices into json automatically
- build an etl pipeline for unstructured documents
- deploy document extraction as an api
- process bank statements without templates
- add document extraction to ai agents via mcp
- reduce llm token costs for document parsing
- automate accounts payable document processing
When to choose
- you need production-grade extraction from varied document formats without training or templates
- you want API or ETL deployment of document extraction workflows
- hallucination control matters - dual-LLM consensus validation is valuable
- you need connectors to file systems and databases for bulk document processing
When to avoid
- you only need simple OCR text extraction without structured output
- you want a lightweight library to embed in code rather than a platform to deploy
- AGPL-3.0 licensing is incompatible with your usage
- your documents are already structured (CSV, JSON) and need no LLM parsing
Facets
application · maturity active
ocr etl rag prompt-engineering mcp llm-inference data-science api-framework self-hosted artificial-intelligence large-language-models pdf developer-tools python self-hosted cloud intelligent-document-processing idp document-ai structured-output json-extraction llm-challenge prompt-studio etl-pipelines agentic-ai pdf-extraction data-engineering automation docker web-server
6 sources
- readme: https://github.com/Zipstack/unstract · fetched 2026-08-28 · 502c7b6ca6c8
- homepage: https://unstract.com · fetched 2026-08-29 · 7523b9e8d5cb
- site_page: https://docs.unstract.com/ · fetched 2026-08-29 · 3b8e138d44cc
- site_page: https://unstract.com/pricing · fetched 2026-08-29 · 19b4c6be2d27
- site_page: https://unstract.com/unstract-editions · fetched 2026-08-29 · c4a68faa5aef
- site_page: https://unstract.com/ai-accounts-payable-procurement-document-processing/contract-pricing-proposal-data-extraction · fetched 2026-08-29 · 27cacae9dbce
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| Zipstack/unstract | main | 88 |
For agents
markdown · JSON · MCP: product_card(name="Zipstack/unstract")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem