Ross ROSS = Recommend OSS · open-source software intelligence for agents

google-research-datasets/natural-questions resource

Natural Questions (NQ) contains real user questions issued to Google search, and answers found from Wikipedia by annotators. NQ is designed for the training and evaluation of automatic question answering systems. observed · 2026-08-28

github.com/google-research-datasets/natural-questions · Python · Apache-2.0 (permissive) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases archived

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2780
  • days_rel: n/a
  • days_push: 1861
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1138 stars · 164 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Natural Questions is a large-scale question answering dataset from Google containing 307k real user search questions paired with Wikipedia pages annotated with long and short answers. The repository provides data utilities, a data browser, and evaluation tooling for training and benchmarking QA systems.

Use cases

  • train a question answering model on real search queries
  • evaluate extractive QA systems against a benchmark
  • get question-answer pairs from Wikipedia pages
  • benchmark long answer and short answer selection models
  • compare my QA system on the Natural Questions leaderboard

When to choose

  • you need a large, realistic QA benchmark based on real user queries
  • you want to compare against published QA baselines and leaderboard results
  • you need Wikipedia-sourced training data with human-annotated answers

When to avoid

  • you need conversational or multi-turn QA data
  • you need small lightweight datasets for quick prototyping
  • you need non-English question answering data

Facets

dataset · maturity maintenance

machine-learning nlp data-science machine-learning python question-answering benchmark wikipedia google-search-queries dataset natural-language-processing search

2 sources

Member repositories

RepositoryRoleHealth v2
google-research-datasets/natural-questionsmain10

For agents

markdown · JSON · MCP: product_card(name="google-research-datasets/natural-questions")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem