srx-2000/spider_collection
python爬虫,目前库存:网易云音乐歌曲爬取,B站视频爬取,知乎问答爬取,壁纸爬取,xvideos视频爬取,有声书爬取,微博爬虫,安居客信息爬取+数据可视化,哔哩哔哩视频封面提取器,ip代理池封装,知乎百万级用户爬虫+数据分析,github用户爬虫 observed · 2026-08-28
Health v2 · maintenance only
32/100
- Activity 0
- Release rhythm 35
- Longevity 100
Flags: no_releases
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 2402
- days_rel: n/a
- days_push: 863
- n_releases_24m: 0
Adoption not part of the score
1646 stars · 239 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
A collection of Python web crawler scripts targeting sites like Bilibili, Zhihu, Weibo, NetEase Music, GitHub, and Anjuke, built with requests, Scrapy, parsel, and BeautifulSoup. It also includes a re-wrapped IP proxy pool and simple data analysis/visualization examples using pandas and matplotlib.
Use cases
- download bilibili videos or extract video covers
- scrape zhihu answers and user profiles for analysis
- collect weibo user info
- download netease cloud music playlists
- scrape rental listings from anjuke and visualize them
- learn how to build python crawlers with scrapy and requests
- set up a wrapped ip proxy pool for anti-scraping
When to choose
- you want ready-made example crawlers for popular Chinese websites
- you are learning web scraping techniques like threading, proxy pools, and scrapy
- you need a starting point to adapt a crawler for bilibili, zhihu, or weibo
When to avoid
- you need a production-grade, maintained scraping framework
- you need a general-purpose crawler with a stable API
- the target site has changed its layout and the script is outdated (some spiders are marked deprecated)
Facets
library · maturity active
web-scraping data-visualization caching developer-tools crawlers data-science social-media developer-tools python cross-platform scrapy requests proxy-pool bilibili zhihu weibo netease-music multi-spider-collection learning-project automation
1 source
- readme: https://github.com/srx-2000/spider_collection · fetched 2026-08-28 · 5b598cff5b9f
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| srx-2000/spider_collection | main | 32 |
For agents
markdown · JSON · MCP: product_card(name="srx-2000/spider_collection")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem