# baidu/DuReader

Baseline Systems of DuReader Dataset

Repository: https://github.com/baidu/DuReader
Canonical: https://ross.abutalabs.com/products/dureader
Homepage: http://ai.baidu.com/broad/subordinate?dataset=dureader
Language: Python
License Family: other
Last push: 2022-05-26T09:30:46+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 3219, "days_push": 1560, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1178, forks 306 (observed 2026-08-28T04:03:52.900497+00:00)

## What it is
DuReader is a collection of Chinese machine reading comprehension and question answering benchmark datasets, including MRC, passage retrieval, DocVQA, robustness, and question matching variants. The repository also provides baseline systems and models such as KT-NET and D-NET for these benchmarks.

## Use cases
- evaluate machine reading comprehension models on Chinese question answering
- benchmark passage retrieval systems on a large-scale Chinese dataset
- test robustness of question matching models against linguistic perturbations
- train and evaluate open-domain document visual question answering models
- compare MRC model generalization across datasets
- download Chinese QA benchmark datasets for research

## When to choose
- you need Chinese-language QA or MRC benchmark datasets with leaderboards
- you want baseline models for reading comprehension research
- you are evaluating model robustness, generalization, or opinion polarity judgment in QA

## When to avoid
- you need English-language QA benchmarks only
- you want a production-ready QA system rather than research datasets and baselines
- you need actively maintained code with a clear license

## Facets
- artifact type: dataset
- maturity: maintenance
- function: machine-learning, nlp, search-engine, benchmarking
- domain: machine-learning, artificial-intelligence
- platform: python
- tags: question-answering, machine-reading-comprehension, chinese-nlp, docvqa, passage-retrieval, baseline-models, natural-language-processing, search

## Member repositories
- baidu/DuReader (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:52.900497+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:26:05.834693+00:00, confidence not recorded.
  - readme: https://github.com/baidu/DuReader (fetched 2026-08-28T04:03:52.900497+00:00, sha a32bb21627fa)
  - homepage: http://ai.baidu.com/broad/subordinate?dataset=dureader (fetched 2026-08-29T12:32:57.763539+00:00, sha 379c350ffcca)
- Data as of 2026-08-30T08:39:29.467469+00:00.
