xishandong/crawlProject resource
python爬虫项目合集,从基础到js逆向,包含基础篇、自动化篇、进阶篇以及验证码篇。案例涵盖各大网站(xhs douyin weibo ins boss job,jd...),你将会学到有关爬虫以及反爬虫、自动化和验证码的各方面知识 observed · 2026-08-28
Health v2 · maintenance only
28/100
- Activity 0
- Release rhythm 35
- Longevity 81
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 1141
- days_rel: n/a
- days_push: 709
- n_releases_24m: 0
Adoption not part of the score
1769 stars · 346 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A collection of hands-on Python web scraping practice projects ranging from beginner requests-based crawlers to JavaScript reverse engineering, browser automation, and captcha solving. It covers real-world cases from sites like Xiaohongshu, Douyin, Weibo, Instagram, Boss Zhipin, and JD, organized into difficulty-rated sections (basics, automation, advanced, captcha).
Use cases
- learn web scraping from scratch with python
- practice bypassing anti-crawler protections
- reverse engineer javascript encryption in website requests
- automate browsers with playwright and selenium
- solve slider and click captchas with ddddocr
- learn scrapy and async crawling
- study real scraping cases for popular chinese websites
When to choose
- you want graded, hands-on crawler exercises from beginner to advanced
- you need examples of js reverse engineering, webpack, browser fingerprinting, and wasm challenges
- you want to learn captcha solving and browser automation in python
When to avoid
- you need a production-ready scraping library or maintained tool rather than practice code
- you plan commercial use - the author explicitly forbids it
- you need guaranteed working examples, since some projects may no longer be reproducible
Facets
learning-resource · maturity active
web-scraping developer-tools workflow-automation crawlers tutorials developer-tools python cross-platform crawler anti-crawler js-reverse-engineering captcha playwright selenium scrapy ddddocr educational-projects automation javascript
1 source
- readme: https://github.com/xishandong/crawlProject · fetched 2026-08-28 · 388b34ea19a9
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| xishandong/crawlProject | main | 28 |
For agents
markdown · JSON · MCP: product_card(name="xishandong/crawlProject")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem