Ross ROSS = Recommend OSS · open-source software intelligence for agents

morphik-org/morphik-core

Open-source multimodal retrieval engine (Morphik Core). By Morphik — AI back office for skilled nursing & senior living (morphik.ai). observed · 2026-08-28

github.com/morphik-org/morphik-core · homepage · Python · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

64/100

  • Activity 94
  • Release rhythm 35
  • Longevity 47

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 660
  • days_rel: n/a
  • days_push: 41
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

3708 stars · 325 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Morphik Core is an open-source, source-available multimodal retrieval engine for building RAG applications over unstructured data like PDFs, videos, and visually rich documents. It provides document ingestion, multimodal embeddings (ColPali), vector storage, retrieval APIs, MCP support, and multi-tenant user/folder scoping via Python/TypeScript SDKs and a REST API.

Use cases

  • build a RAG pipeline over PDFs and visually rich documents
  • search tables and charts inside scanned documents with multimodal embeddings
  • ingest videos and PDFs into a searchable knowledge base
  • add retrieval-augmented generation to an AI application without stitching together OCR, embeddings, and a vector DB
  • build a multi-tenant AI app with per-user and per-folder data isolation
  • expose a document knowledge base to MCP clients like Claude
  • query documents with an LLM and get grounded answers

When to choose

  • you need production RAG over complex, visually rich documents where plain text extraction fails
  • you want an all-in-one ingestion, embedding, storage, and retrieval platform instead of assembling separate tools
  • you need multi-user/folder scoping for multi-tenant applications
  • you want built-in MCP support to plug your knowledge base into AI agents

When to avoid

  • you only need a lightweight vector store and already have your own ingestion and embedding pipeline
  • you require a permissively licensed OSS dependency (license is source-available, not standard OSI)
  • you need fully offline air-gapped operation but rely on the hosted Morphik cloud API
  • your data is purely structured/tabular relational data better served by a SQL database

Facets

service · maturity active

rag vector-database search-engine nlp mcp llm-inference pdf etl artificial-intelligence large-language-models databases pdf python self-hosted cross-platform multimodal-retrieval colpali cache-augmented-generation unstructured-data knowledge-base document-ingestion source-available multitenancy retrieval-augmented-generation search natural-language-processing ai-agents docker web-server

9 sources

Member repositories

RepositoryRoleHealth v2
morphik-org/morphik-coremain64

For agents

markdown · JSON · MCP: product_card(name="morphik-org/morphik-core")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem