# juand-r/entity-recognition-datasets

A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.

Repository: https://github.com/juand-r/entity-recognition-datasets
Canonical: https://ross.abutalabs.com/products/entity-recognition-datasets
Language: Python
License: MIT
License Family: permissive
Topics: entity-extraction, named-entity-recognition, ner, datasets, entity-recognition, nlp-resources, nlp, corpora, natural-language-processing, annotations
Last push: 2026-07-02T04:27:10+00:00

## Health v2 (maintenance only)
Score: 73/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 90, release rhythm 35, longevity 100
- inputs: {"age_days": 2923, "days_push": 62, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1574, forks 245 (observed 2026-08-28T04:05:05.772324+00:00)

## What it is
A curated collection of annotated corpora for named entity recognition (NER) and entity recognition tasks across multiple languages, domains, and entity types. It includes conversion code to the CoNLL 2003 format and links to datasets that cannot be redistributed due to licensing.

## Use cases
- find NER training datasets for English
- get entity recognition corpora in other languages
- convert NER datasets to CoNLL 2003 format
- benchmark named entity recognition models on multiple corpora
- find domain-specific NER data like Twitter or Wikipedia text
- collect annotated data for sequence labeling research

## When to choose
- you need a catalog of NER corpora with licensing and availability info
- you want datasets pre-converted or convertible to CoNLL format
- you are comparing NER models across multiple benchmarks

## When to avoid
- you need recently published NER datasets after 2020, as the list is no longer actively updated
- you need the actual data for LDC-licensed corpora, which must be purchased separately
- you need a hosted dataset API rather than links and conversion scripts

## Facets
- artifact type: dataset
- maturity: maintenance
- function: nlp, data-science
- domain: machine-learning, data-science
- platform: python, cross-platform
- tags: named-entity-recognition, ner, corpora, entity-extraction, conll-format, dataset-collection, natural-language-processing

## Member repositories
- juand-r/entity-recognition-datasets (main) score 73

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:05.772324+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:57:39.353391+00:00, confidence not recorded.
  - readme: https://github.com/juand-r/entity-recognition-datasets (fetched 2026-08-28T04:05:05.772324+00:00, sha 111cb21c0fa5)
- Data as of 2026-08-30T08:39:29.467469+00:00.
