Ross ROSS = Recommend OSS · open-source software intelligence for agents

zai-org/GLM-OCR

GLM-OCR: Accurate × Fast × Comprehensive observed · 2026-08-28

github.com/zai-org/GLM-OCR · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

65/100

  • Activity 78
  • Release rhythm 78
  • Longevity 15
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 2.5
  • age_days: 212
  • days_rel: 147
  • days_push: 134
  • n_releases_24m: 5

Full methodology

Adoption not part of the score

7366 stars · 664 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

GLM-OCR is an open-source 0.9B-parameter multimodal OCR model built on the GLM-V encoder-decoder architecture for complex document understanding, including tables, formulas, and code-heavy layouts. It ships with a Python SDK, CLI, and inference toolchain supporting deployment via vLLM, SGLang, and Ollama.

Use cases

  • extract text from scanned pdf documents
  • recognize tables in images
  • convert formulas in papers to latex
  • parse complex document layouts into structured text
  • run ocr at scale on a gpu server
  • digitize receipts and sealed documents

When to choose

  • you need state-of-the-art document OCR with layout analysis
  • you want a small efficient model deployable with vLLM or Ollama
  • your documents contain tables, formulas, seals, or code

When to avoid

  • you only need simple plain-text OCR without layout understanding
  • you have no GPU and cannot use the hosted API
  • you need handwriting recognition in many languages

Facets

library · maturity active

ocr machine-learning llm-inference image-processing pdf computer-vision artificial-intelligence pdf deep-learning python cross-platform cli multimodal document-understanding vision-language-model table-recognition formula-recognition vllm sglang ollama layout-analysis natural-language-processing gpu

1 source

Member repositories

RepositoryRoleHealth v2
zai-org/GLM-OCRmain65

For agents

markdown · JSON · MCP: product_card(name="zai-org/GLM-OCR")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem