# xishandong/crawlProject

python爬虫项目合集，从基础到js逆向，包含基础篇、自动化篇、进阶篇以及验证码篇。案例涵盖各大网站(xhs douyin weibo ins boss job，jd...)，你将会学到有关爬虫以及反爬虫、自动化和验证码的各方面知识

Repository: https://github.com/xishandong/crawlProject
Canonical: https://ross.abutalabs.com/products/crawlproject
Language: JavaScript
License Family: other
Topics: playwright, python, python-crawler, reverse-engineering, captcha, ddddocr, javascript
Last push: 2024-09-23T11:27:25+00:00

## Health v2 (maintenance only)
Score: 28/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 81
- inputs: {"age_days": 1141, "days_push": 709, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1769, forks 346 (observed 2026-08-28T04:05:33.845520+00:00)

## What it is
A collection of hands-on Python web scraping practice projects ranging from beginner requests-based crawlers to JavaScript reverse engineering, browser automation, and captcha solving. It covers real-world cases from sites like Xiaohongshu, Douyin, Weibo, Instagram, Boss Zhipin, and JD, organized into difficulty-rated sections (basics, automation, advanced, captcha).

## Use cases
- learn web scraping from scratch with python
- practice bypassing anti-crawler protections
- reverse engineer javascript encryption in website requests
- automate browsers with playwright and selenium
- solve slider and click captchas with ddddocr
- learn scrapy and async crawling
- study real scraping cases for popular chinese websites

## When to choose
- you want graded, hands-on crawler exercises from beginner to advanced
- you need examples of js reverse engineering, webpack, browser fingerprinting, and wasm challenges
- you want to learn captcha solving and browser automation in python

## When to avoid
- you need a production-ready scraping library or maintained tool rather than practice code
- you plan commercial use - the author explicitly forbids it
- you need guaranteed working examples, since some projects may no longer be reproducible

## Facets
- artifact type: learning-resource
- maturity: active
- function: web-scraping, developer-tools, workflow-automation
- domain: crawlers, tutorials, developer-tools
- platform: python, cross-platform
- tags: crawler, anti-crawler, js-reverse-engineering, captcha, playwright, selenium, scrapy, ddddocr, educational-projects, automation, javascript

## Member repositories
- xishandong/crawlProject (main) score 28

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:33.845520+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:26:08.420329+00:00, confidence not recorded.
  - readme: https://github.com/xishandong/crawlProject (fetched 2026-08-28T04:05:33.845520+00:00, sha 388b34ea19a9)
- Data as of 2026-08-30T08:39:29.467469+00:00.
