webrecorder/browsertrix-crawler
Run a high-fidelity browser-based web archiving crawler in a single Docker container observed · 2026-08-28
Health v2 · maintenance only
99/100
- Activity 99
- Release rhythm 98
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: 9.0
- age_days: 2130
- days_rel: 13
- days_push: 7
- n_releases_24m: 55
Adoption not part of the score
1120 stars · 149 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Browsertrix Crawler is a standalone browser-based high-fidelity web crawling system that runs in a single Docker container. It uses Puppeteer to control Brave Browser windows and captures pages via the Chrome Devtools Protocol, producing WARC/WACZ web archives.
Use cases
- crawl websites into WARC/WACZ archives
- run a high-fidelity browser-based web crawler in Docker
- archive dynamic JavaScript-heavy pages with screenshots and autoscroll
- crawl sites behind logins using saved browser profiles
- run QA analysis on replayed web crawls
- capture websites for offline access or digital preservation
When to choose
- you need browser-rendered, high-fidelity captures of modern JavaScript-heavy sites
- you want standards-compliant WARC/WACZ output for web archiving
- you need a self-contained crawler that runs in a single Docker container
- you need login-based crawling, custom behaviors, or per-seed scoping rules
When to avoid
- you only need lightweight raw HTML fetching without browser rendering
- you cannot run Docker containers in your environment
- you need a distributed crawling cluster rather than a single-container crawler
Facets
cli-tool · maturity active
web-scraping cli developer-tools crawlers web-development developer-tools cli windows web-archiving warc wacz puppeteer crawler headless-browser docker automation linux macos
3 sources
- readme: https://github.com/webrecorder/browsertrix-crawler · fetched 2026-08-28 · 0ba57ddc578a
- homepage: https://crawler.docs.browsertrix.com · fetched 2026-08-29 · 96d3865ddfa0
- site_page: https://crawler.docs.browsertrix.com/develop/docs · fetched 2026-08-29 · 1781ee02a743
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| webrecorder/browsertrix-crawler | main | 99 |
For agents
markdown · JSON · MCP: product_card(name="webrecorder/browsertrix-crawler")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem