ibm-aur-nlp/PubLayNet resource
None observed · 2026-08-28
Health v2 · maintenance only
46/100
- Activity 30
- Release rhythm 35
- Longevity 100
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 2680
- days_rel: n/a
- days_push: 421
- n_releases_24m: 0
Adoption not part of the score
1060 stars · 167 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
PubLayNet is a large annotated dataset of over 360k document images from PubMed Central, with bounding boxes and polygonal segmentations for layout elements like text, titles, lists, tables, and figures. The repository also hosts pre-trained Faster-RCNN and Mask-RCNN models and ICDAR 2021 Scientific Literature Parsing competition materials.
Use cases
- train a document layout detection model
- segment text, tables, and figures in scientific paper images
- benchmark object detection models on document images
- extract structure from PDF pages of research articles
- download pre-trained Mask-RCNN for document analysis
- participate in scientific literature parsing competitions
When to choose
- you need large-scale labeled data for document layout analysis
- you are training or evaluating object detection models on scientific documents
- you want pre-trained models for parsing PubMed-style paper layouts
When to avoid
- you need table structure recognition rather than layout detection (see PubTabNet)
- you need annotated documents outside the scientific/medical domain
- you need a ready-made application rather than a dataset and models
Facets
dataset · maturity maintenance
computer-vision ocr machine-learning data-science computer-vision machine-learning python cross-platform document-layout-analysis object-detection image-segmentation scientific-documents pubmed faster-rcnn mask-rcnn icdar natural-language-processing datasets
1 source
- readme: https://github.com/ibm-aur-nlp/PubLayNet · fetched 2026-08-28 · 58969915b228
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| ibm-aur-nlp/PubLayNet | main | 46 |
For agents
markdown · JSON · MCP: product_card(name="ibm-aur-nlp/PubLayNet")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem