# tb0hdan/domains

World’s single largest Internet domains dataset

Repository: https://github.com/tb0hdan/domains
Canonical: https://ross.abutalabs.com/products/domains
Homepage: https://domainsproject.org
Language: JavaScript
License: BSD-3-Clause
License Family: permissive
Topics: yacy, scrapy, dataset, search-engines, internet-domains, colly
Last push: 2026-05-03T00:27:11+00:00

## Health v2 (maintenance only)
Score: 68/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 80, release rhythm 35, longevity 100
- inputs: {"age_days": 2425, "days_push": 123, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1153, forks 179 (observed 2026-08-28T04:03:47.561489+00:00)

## What it is
A public dataset of billions of sorted Internet domains, described as the world's largest, built by processing petabytes of crawl and DNS observation data. It is continuously updated with Newly Observed Hostnames and offered under a BSD-3-Clause license, with subscription tiers for extended data like historical forward DNS records.

## Use cases
- get a list of all registered internet domains
- build a domain intelligence or threat feed
- research DNS censorship or domain registry operations
- seed a web crawler with known domains
- check whether a domain exists in a large corpus
- analyze TLD coverage and domain name trends
- feed domain data into a search engine index

## When to choose
- you need a massive, regularly updated corpus of internet domains for research, security, or search
- you want freely licensed domain data without scraping it yourself
- you need first-seen hostname feeds for threat intelligence

## When to avoid
- you need per-domain ownership or WHOIS details rather than just names
- you need historical DNS records without a subscription
- your use case requires only a small curated domain list

## Facets
- artifact type: dataset
- maturity: active
- function: search-engine, data-science, web-scraping, analytics
- domain: security, big-data, networking, osint
- platform: cross-platform, cli
- tags: internet-domains, dns, passive-dns, newly-observed-hostnames, tld-coverage, domain-list, crawler, search

## Member repositories
- tb0hdan/domains (main) score 68

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:47.561489+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:33:29.911975+00:00, confidence not recorded.
  - readme: https://github.com/tb0hdan/domains (fetched 2026-08-28T04:03:47.561489+00:00, sha 8aff43abe2fd)
  - homepage: https://domainsproject.org (fetched 2026-08-29T12:38:24.492345+00:00, sha 978b01d66e32)
- Data as of 2026-08-30T08:39:29.467469+00:00.
