# unsplash/datasets

🎁  7,400,000+ Unsplash images made available for research and machine learning

Repository: https://github.com/unsplash/datasets
Canonical: https://ross.abutalabs.com/products/unsplash-datasets
Homepage: https://unsplash.com/data
Language: Jupyter Notebook
License Family: other
Topics: dataset, images, unsplash, machine-learning, research, data, search-engine, keywords, photos, semantics
Last push: 2026-06-26T15:30:57+00:00

## Health v2 (maintenance only)
Score: 80/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 89, release rhythm 58, longevity 100
- inputs: {"age_days": 2262, "days_push": 68, "days_rel": 68, "gap_med": 217.5, "n_releases_24m": 3}
- flags: no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2777, forks 142 (observed 2026-08-28T04:07:21.524725+00:00)

## What it is
The Unsplash Dataset is a collection of over 7.4 million high-quality photos with rich metadata including keywords, search queries, and camera information, released by Unsplash for research and machine learning. It comes in two tiers: a Lite dataset (~25k photos, commercially usable) and a Full dataset (non-commercial, access-request required).

## Use cases
- train image classification models on labeled photos
- research visual search and semantic image retrieval
- build image search engines with real search query data
- analyze photography trends from camera and lens metadata
- train multimodal image-text models
- benchmark image quality and popularity prediction models

## When to choose
- you need a large, openly licensed image dataset with real-world search and keyword metadata
- you want commercially usable image data for prototyping (Lite tier)
- you are researching semantic visual search beyond object detection

## When to avoid
- you need commercial usage rights for the full 7.4M photo dataset
- you need user-level or personally identifiable data (never included)
- you need video or audio data
- you require a formal open-source license (the repo has none; usage is governed by separate terms)

## Facets
- artifact type: dataset
- maturity: active
- function: machine-learning, search-engine, image-processing, data-science
- domain: machine-learning, computer-vision, image-processing, data-science, photography
- platform: python, cross-platform
- tags: image-dataset, unsplash, photo-metadata, visual-search, keywords, research-dataset, computer-vision-training, search

## Member repositories
- unsplash/datasets (main) score 80

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:21.524725+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T08:16:31.756338+00:00, confidence not recorded.
  - readme: https://github.com/unsplash/datasets (fetched 2026-08-28T04:07:21.524725+00:00, sha 88a0447e0278)
  - homepage: https://unsplash.com/data (fetched 2026-08-29T09:55:54.022070+00:00, sha 3acce05d9b37)
  - site_page: https://unsplash.com/documentation (fetched 2026-08-29T09:55:54.031735+00:00, sha f7915740cad3)
  - site_page: https://unsplash.com/documentation/changelog (fetched 2026-08-29T09:55:54.036900+00:00, sha 66513297101f)
- Data as of 2026-08-30T08:39:29.467469+00:00.
