# google-research-datasets/wit

WIT (Wikipedia-based Image Text) Dataset is a large multimodal multilingual dataset comprising 37M+ image-text sets with 11M+ unique images across 100+ languages.

Repository: https://github.com/google-research-datasets/wit
Canonical: https://ross.abutalabs.com/products/wit
Homepage: https://github.com/google-research-datasets/wit
License: NOASSERTION
License Family: other
Topics: nlp, machine-learning, wikipedia, multimodal, multilingual, cc-by-sa-3
Archived: true
Last push: 2024-09-27T20:55:42+00:00

## Health v2 (maintenance only)
Score: 10/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2016, "days_push": 705, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, archived, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1113, forks 47 (observed 2026-08-28T04:03:38.116061+00:00)

## What it is
WIT (Wikipedia-based Image Text) is a large multimodal multilingual dataset of 37.6 million image-text pairs with 11.5 million unique images across 108 Wikipedia languages. It is intended for pretraining and evaluating multimodal machine learning models such as vision-language models.

## Use cases
- pretrain a vision-language model on image-text pairs
- find a large multilingual image captioning dataset
- train a CLIP-style multimodal model
- evaluate cross-lingual image-text retrieval
- get image-text training data covering 100+ languages
- benchmark multimodal models on real-world entities

## When to choose
- you need a massive, publicly available image-text corpus for multimodal pretraining
- you need multilingual coverage across 108 languages
- you want page-level metadata and contextual information alongside image-text pairs

## When to avoid
- you need a small, curated dataset for quick experiments
- you require commercial licensing beyond CC BY-SA 3.0 terms
- you need actively maintained tooling rather than a static dataset release

## Facets
- artifact type: dataset
- maturity: maintenance
- function: machine-learning, nlp, image-processing
- domain: machine-learning, computer-vision, large-language-models
- platform: python, cross-platform
- tags: multimodal, multilingual, image-text, wikipedia, pretraining, vision-language, dataset, natural-language-processing

## Member repositories
- google-research-datasets/wit (main) score 10

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:38.116061+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:42:36.698234+00:00, confidence not recorded.
  - readme: https://github.com/google-research-datasets/wit (fetched 2026-08-28T04:03:38.116061+00:00, sha 17f39a59724d)
  - homepage: https://github.com/google-research-datasets/wit (fetched 2026-08-29T12:46:37.289644+00:00, sha ce69c99827e6)
- Data as of 2026-08-30T08:39:29.467469+00:00.
