# Kr1s77/Python-crawler-tutorial-starts-from-zero

python爬虫教程，带你从零到一，包含js逆向，selenium, tesseract OCR识别,mongodb的使用，以及scrapy框架

Repository: https://github.com/Kr1s77/Python-crawler-tutorial-starts-from-zero
Canonical: https://ross.abutalabs.com/products/python-crawler-tutorial-starts-from-zero
Language: Python
License Family: other
Last push: 2020-12-02T03:01:29+00:00

## Health v2 (maintenance only)
Score: 32/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 35, longevity 100
- inputs: {"age_days": 2715, "days_push": 2100, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 4613, forks 761 (observed 2026-08-28T04:08:55.073111+00:00)

## What it is
A Chinese-language tutorial repository teaching Python web scraping from zero, covering requests, data extraction, JS reverse engineering, Selenium, Tesseract OCR, MongoDB, and the Scrapy framework. It combines written lessons with practical example crawlers for sites like Douban and Baidu.

## Use cases
- learn python web scraping from scratch
- understand how to reverse engineer javascript for crawlers
- scrape websites with selenium
- recognize captcha images with tesseract ocr
- store scraped data in mongodb
- learn the scrapy framework
- extract  and regex data from web pages

## When to choose
- you are a beginner wanting a structured Chinese-language crawler tutorial
- you want hands-on examples covering requests, selenium, scrapy, and OCR in one place

## When to avoid
- you need an actively maintained production scraping library
- you need English-language documentation
- you need a tool rather than learning material

## Facets
- artifact type: learning-resource
- maturity: maintenance
- function: web-scraping, ocr, parser, developer-tools
- domain: crawlers, tutorials, developer-tools
- platform: python, cross-platform
- tags: python-crawler, js-reverse-engineering, selenium, scrapy, mongodb, tesseract-ocr, chinese-language, natural-language-processing

## Member repositories
- Kr1s77/Python-crawler-tutorial-starts-from-zero (main) score 32

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:55.073111+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:19:41.843867+00:00, confidence not recorded.
  - readme: https://github.com/Kr1s77/Python-crawler-tutorial-starts-from-zero (fetched 2026-08-28T04:08:55.073111+00:00, sha c17ff79f1b4e)
- Data as of 2026-08-30T08:39:29.467469+00:00.
