Ross ROSS = Recommend OSS · open-source software intelligence for agents

grobidOrg/grobid

A machine learning software for extracting information from scholarly documents observed · 2026-08-28

github.com/grobidOrg/grobid · homepage · Java · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

87/100

  • Activity 99
  • Release rhythm 64
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: 239
  • age_days: 5102
  • days_rel: 29
  • days_push: 8
  • n_releases_24m: 4

Full methodology

Adoption not part of the score

5105 stars · 567 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

GROBID is a Java machine learning library that extracts, parses, and re-structures raw documents such as PDFs into structured XML/TEI, focused on technical and scientific publications. It provides header extraction, reference parsing, citation context resolution, and full-text document segmentation, with a REST API and Docker deployment.

Use cases

  • extract metadata (title, authors, abstract) from scientific PDFs
  • parse bibliographic references from scholarly articles
  • convert PDF papers into structured TEI XML
  • resolve citation contexts to full references
  • segment and structure full text of academic PDFs
  • batch process large corpora of scientific publications

When to choose

  • you need high-accuracy extraction of bibliographic data from scholarly PDFs
  • you want structured XML/TEI output from unstructured scientific documents
  • you need a self-hosted, Docker-deployable PDF parsing service
  • you are building scholarly search, citation analysis, or literature review tooling

When to avoid

  • you only need simple text extraction without structure (pdftotext suffices)
  • your documents are not technical/scientific publications
  • you cannot run a JVM-based service or ML models
  • you need extraction from scanned PDFs without OCR preprocessing

Facets

library · maturity active

machine-learning nlp pdf parser ocr machine-learning pdf developer-tools jvm cross-platform pdf-parsing tei-xml scholarly-documents bibliographic-references metadata-extraction deep-learning crf transformers natural-language-processing docker web-server

1 source

Member repositories

RepositoryRoleHealth v2
grobidOrg/grobidmain87

For agents

markdown · JSON · MCP: product_card(name="grobidOrg/grobid")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem