Ross ROSS = Recommend OSS · open-source software intelligence for agents

datalab-to/surya

OCR, layout analysis, reading order, table recognition in 90+ languages observed · 2026-08-28

github.com/datalab-to/surya · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

86/100

  • Activity 98
  • Release rhythm 81
  • Longevity 69
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 3
  • age_days: 966
  • days_rel: 44
  • days_push: 12
  • n_releases_24m: 62

Full methodology

Adoption not part of the score

21318 stars · 1533 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Surya is a 650M parameter OCR toolkit from Datalab providing state-of-the-art text recognition, layout analysis, reading order detection, and table recognition in 90+ languages. It ships as a Python library with GPU acceleration and is part of Datalab's document intelligence model suite.

Use cases

  • extract text from scanned pdfs
  • ocr documents in multiple languages
  • detect layout regions and reading order in documents
  • recognize tables from images
  • convert scanned documents to markdown
  • run ocr on gpu at high throughput

When to choose

  • you need multilingual ocr across 90+ languages
  • you need layout analysis and table recognition alongside text extraction
  • you want a self-hosted, Apache-2.0 licensed ocr pipeline with gpu speed
  • you are building document processing or rag ingestion pipelines

When to avoid

  • you only need simple English-only ocr with a lightweight dependency
  • you have no gpu and need very fast batch processing
  • you need a fully managed api without running models yourself
  • your models must be under a permissive license without usage restrictions (weights use OpenRAIL-M)

Facets

library · maturity active

ocr image-processing machine-learning deep-learning pdf computer-vision pdf machine-learning files python cross-platform document-intelligence layout-analysis table-recognition reading-order multilingual text-detection natural-language-processing gpu

2 sources

Member repositories

RepositoryRoleHealth v2
datalab-to/suryamain86

For agents

markdown · JSON · MCP: product_card(name="datalab-to/surya")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem