Ross ROSS = Recommend OSS · open-source software intelligence for agents

ibm-aur-nlp/PubLayNet resource

None observed · 2026-08-28

github.com/ibm-aur-nlp/PubLayNet · Jupyter Notebook · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

46/100

  • Activity 30
  • Release rhythm 35
  • Longevity 100

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2680
  • days_rel: n/a
  • days_push: 421
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1060 stars · 167 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

PubLayNet is a large annotated dataset of over 360k document images from PubMed Central, with bounding boxes and polygonal segmentations for layout elements like text, titles, lists, tables, and figures. The repository also hosts pre-trained Faster-RCNN and Mask-RCNN models and ICDAR 2021 Scientific Literature Parsing competition materials.

Use cases

  • train a document layout detection model
  • segment text, tables, and figures in scientific paper images
  • benchmark object detection models on document images
  • extract structure from PDF pages of research articles
  • download pre-trained Mask-RCNN for document analysis
  • participate in scientific literature parsing competitions

When to choose

  • you need large-scale labeled data for document layout analysis
  • you are training or evaluating object detection models on scientific documents
  • you want pre-trained models for parsing PubMed-style paper layouts

When to avoid

  • you need table structure recognition rather than layout detection (see PubTabNet)
  • you need annotated documents outside the scientific/medical domain
  • you need a ready-made application rather than a dataset and models

Facets

dataset · maturity maintenance

computer-vision ocr machine-learning data-science computer-vision machine-learning python cross-platform document-layout-analysis object-detection image-segmentation scientific-documents pubmed faster-rcnn mask-rcnn icdar natural-language-processing datasets

1 source

Member repositories

RepositoryRoleHealth v2
ibm-aur-nlp/PubLayNetmain46

For agents

markdown · JSON · MCP: product_card(name="ibm-aur-nlp/PubLayNet")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem