grobidOrg/grobid
A machine learning software for extracting information from scholarly documents observed · 2026-08-28
Health v2 · maintenance only
87/100
- Activity 99
- Release rhythm 64
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 239
- age_days: 5102
- days_rel: 29
- days_push: 8
- n_releases_24m: 4
Adoption not part of the score
5105 stars · 567 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
GROBID is a Java machine learning library that extracts, parses, and re-structures raw documents such as PDFs into structured XML/TEI, focused on technical and scientific publications. It provides header extraction, reference parsing, citation context resolution, and full-text document segmentation, with a REST API and Docker deployment.
Use cases
- extract metadata (title, authors, abstract) from scientific PDFs
- parse bibliographic references from scholarly articles
- convert PDF papers into structured TEI XML
- resolve citation contexts to full references
- segment and structure full text of academic PDFs
- batch process large corpora of scientific publications
When to choose
- you need high-accuracy extraction of bibliographic data from scholarly PDFs
- you want structured XML/TEI output from unstructured scientific documents
- you need a self-hosted, Docker-deployable PDF parsing service
- you are building scholarly search, citation analysis, or literature review tooling
When to avoid
- you only need simple text extraction without structure (pdftotext suffices)
- your documents are not technical/scientific publications
- you cannot run a JVM-based service or ML models
- you need extraction from scanned PDFs without OCR preprocessing
Facets
library · maturity active
machine-learning nlp pdf parser ocr machine-learning pdf developer-tools jvm cross-platform pdf-parsing tei-xml scholarly-documents bibliographic-references metadata-extraction deep-learning crf transformers natural-language-processing docker web-server
1 source
- readme: https://github.com/grobidOrg/grobid · fetched 2026-08-28 · b6e85af9a6d2
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| grobidOrg/grobid | main | 87 |
For agents
markdown · JSON · MCP: product_card(name="grobidOrg/grobid")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem