# lorien/awesome-web-scraping

List of libraries, tools and APIs for web scraping and data processing.

Repository: https://github.com/lorien/awesome-web-scraping
Canonical: https://ross.abutalabs.com/products/awesome-web-scraping
License: NOASSERTION
License Family: other
Topics: web-scraping, captcha-bypass, captcha-recaptcha, crawling, crawling-framework, crawling-python, crawling-tool, scraping, scraping-framework, scraping-python, scraping-tool, webscraping, crawler, spider
Last push: 2026-08-26T23:36:29+00:00

## Health v2 (maintenance only)
Score: 77/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 35, longevity 100
- inputs: {"age_days": 4039, "days_push": 7, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 8133, forks 930 (observed 2026-08-28T04:10:13.823947+00:00)

## What it is
A curated awesome-list of web scraping libraries, tools, APIs, and manuals across Python, PHP, Ruby, JavaScript, Go, and CLI tools. It also links to captcha solving services, proxy marketplaces, headless browser lists, and community discussion groups.

## Use cases
- find a web scraping library for python
- discover crawling frameworks for go or ruby
- learn web scraping from tutorials and books
- find captcha solving services
- compare headless browsers for scraping
- find cli tools for scraping websites

## When to choose
- you want a curated starting point for scraping tools in a specific language
- you need links to manuals, proxies, and captcha services in one place

## When to avoid
- you need a working scraping tool rather than a list of links
- you need guaranteed licensing or maintained code, since it is a list with a nonstandard license

## Facets
- artifact type: learning-resource
- maturity: active
- function: web-scraping, developer-tools
- domain: crawlers, web-development, awesome-lists
- platform: cross-platform
- tags: awesome-list, crawling, scraping, captcha-solving, proxies, curated-list

## Member repositories
- lorien/awesome-web-scraping (main) score 77

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:10:13.823947+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T17:29:51.317740+00:00, confidence not recorded.
  - readme: https://github.com/lorien/awesome-web-scraping (fetched 2026-08-28T04:10:13.823947+00:00, sha aff108e1987f)
- Data as of 2026-08-30T08:39:29.467469+00:00.
