Ross ROSS = Recommend OSS · open-source software intelligence for agents

Unstructured-IO/unstructured

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding. observed · 2026-08-28

github.com/Unstructured-IO/unstructured · homepage · HTML · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

95/100

  • Activity 99
  • Release rhythm 86
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 4.0
  • age_days: 1437
  • days_rel: 12
  • days_push: 9
  • n_releases_24m: 89

Full methodology

Adoption not part of the score

15349 stars · 1307 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Unstructured is an open-source ETL library and platform for converting complex documents (PDF, DOCX, HTML, images, and 65+ file types) into clean, structured JSON for language models. It provides partitioning, chunking, OCR, and enrichment pipelines that feed RAG and agentic AI workflows.

Use cases

  • parse resumes from pdfs
  • convert pdf documents to text for llm ingestion
  • extract structured from docx and html files
  • build a rag pipeline over a document corpus
  • ocr scanned documents into machine-readable text
  • chunk documents for vector database embedding
  • preprocess documents for chatbot knowledge bases

When to choose

  • you need to turn messy documents of many formats into LLM-ready structured data
  • you want an open-source document parsing pipeline with optional hosted platform
  • you're building RAG or agentic AI applications that ingest PDFs, Office files, or HTML

When to avoid

  • you only need simple text extraction from plain text or well-structured markdown
  • you need a lightweight tool without heavy ML dependencies
  • your use case is real-time streaming transformation unrelated to documents

Facets

library · maturity active

etl ocr pdf parser rag nlp data-science pdf large-language-models python self-hosted cloud document-parsing pdf-to-text document-etl chunking preprocessing langchain llm-preparation natural-language-processing data-engineering retrieval-augmented-generation docker

4 sources

Member repositories

RepositoryRoleHealth v2
Unstructured-IO/unstructuredmain95

For agents

markdown · JSON · MCP: product_card(name="Unstructured-IO/unstructured")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem