tjmlabs/ColiVara
Colivara is a suite of services that allows you to store, search, and retrieve documents based on their visual embedding. ColiVara has state of the art retrieval performance on both text and visual documents. using vision models instead of chunking and text-processing for documents. No OCR, no text extraction, no broken tables, or missing images. observed · 2026-08-28
Health v2 · maintenance only
65/100
- Activity 91
- Release rhythm 40
- Longevity 50
Flags: no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 1
- age_days: 712
- days_rel: 573
- days_push: 56
- n_releases_24m: 28
Adoption not part of the score
1486 stars · 117 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
ColiVara is a hosted (and self-hostable) retrieval API that stores, searches, and retrieves documents using visual embeddings generated by vision models like ColQwen2, based on the ColPali paper. It replaces traditional OCR, chunking, and text extraction pipelines, supporting 100+ file formats with Python and TypeScript SDKs.
Use cases
- build a RAG application that understands tables, charts, and diagrams
- search PDFs and documents without OCR or text extraction
- retrieve documents by visual similarity using embeddings
- index complex financial reports and technical diagrams for semantic search
- filter document search results by metadata and collections
- avoid broken tables and missing images in document retrieval
When to choose
- your documents are visually rich with tables, figures, or complex layouts that OCR pipelines break
- you want state-of-the-art document retrieval without building a chunking/embedding pipeline yourself
- you need a simple REST API with Python/TypeScript SDKs for document search
- you want to self-host a ColPali-style visual retrieval service
When to avoid
- you only need plain-text keyword search over simple text documents
- you need full control over chunking and embedding strategies
- you cannot send documents to a third-party API and cannot self-host
- you need a lightweight local library rather than a hosted service
Facets
service · maturity active
rag vector-database search-engine ocr machine-learning sdk api-framework large-language-models artificial-intelligence pdf databases self-hosted python visual-embeddings colpali document-retrieval vision-models multimodal-rag document-search retrieval-augmented-generation search web-server nodejs docker
3 sources
- readme: https://github.com/tjmlabs/ColiVara · fetched 2026-08-28 · 0a1159d8cafe
- homepage: https://colivara.com · fetched 2026-08-29 · 9f29aca94b68
- site_page: https://docs.colivara.com · fetched 2026-08-29 · e84c3004a399
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| tjmlabs/ColiVara | main | 65 |
For agents
markdown · JSON · MCP: product_card(name="tjmlabs/ColiVara")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem