# mittagessen/kraken

OCR engine for all the languages

Repository: https://github.com/mittagessen/kraken
Canonical: https://ross.abutalabs.com/products/mittagessen-kraken
Homepage: http://kraken.re
Language: Python
License: Apache-2.0
License Family: permissive
Topics: ocr, neural-networks, alto-xml, hocr, handwritten-text-recognition, htr, layout-analysis, optical-character-recognition, page-xml
Last push: 2026-08-30T10:55:32+00:00

## Health v2 (maintenance only)
Score: 99/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 100, release rhythm 96, longevity 100
- inputs: {"age_days": 4124, "days_push": 3, "days_rel": 28, "gap_med": 17, "n_releases_24m": 10}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1061, forks 173 (observed 2026-09-01T02:14:05.024554+00:00)

## What it is
kraken is a turn-key OCR/HTR engine built on neural networks, optimized for historical and non-Latin script material. It provides trainable layout analysis, reading order detection, and multi-script character recognition with output in ALTO, PageXML, abbyyXML, and hOCR formats.

## Use cases
- recognize text in historical manuscripts
- OCR for non-Latin scripts like Arabic or Hebrew
- transcribe handwritten documents
- extract text from scanned images into PageXML or ALTO
- binarize and segment scanned page images
- train custom OCR models for rare scripts
- handle right-to-left and bidirectional text in OCR

## When to choose
- you need OCR for historical, damaged, or non-Latin script material
- you want fully trainable recognition and layout analysis models
- you need structured output formats like ALTO, PageXML, or hOCR
- you work with right-to-left, BiDi, or vertical scripts
- you need both printed text and handwritten text recognition

## When to avoid
- you need a plug-and-play OCR for modern Latin-script documents with no training
- you want a GUI-based OCR tool
- you need Windows support
- you want commercial support or prebuilt binaries outside pip

## Facets
- artifact type: library
- maturity: active
- function: ocr, machine-learning, image-processing, nlp
- domain: computer-vision, pdf, files
- platform: python, cli
- tags: handwritten-text-recognition, htr, layout-analysis, pagexml, alto-xml, hocr, historical-documents, non-latin-scripts, neural-networks, natural-language-processing, linux, macos

## Member repositories
- mittagessen/kraken (main) score 99

## Provenance
- Observed fields: from GitHub, fetched 2026-09-01T02:14:05.024554+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:56:58.335527+00:00, confidence not recorded.
  - readme: https://github.com/mittagessen/kraken (fetched 2026-09-01T02:14:05.024554+00:00, sha cc60770dd748)
  - homepage: http://kraken.re (fetched 2026-08-29T12:59:05.427625+00:00, sha 08f06ae12298)
- Data as of 2026-08-30T08:39:29.467469+00:00.
