Ross ROSS = Recommend OSS · open-source software intelligence for agents

srx-2000/spider_collection

python爬虫,目前库存:网易云音乐歌曲爬取,B站视频爬取,知乎问答爬取,壁纸爬取,xvideos视频爬取,有声书爬取,微博爬虫,安居客信息爬取+数据可视化,哔哩哔哩视频封面提取器,ip代理池封装,知乎百万级用户爬虫+数据分析,github用户爬虫 observed · 2026-08-28

github.com/srx-2000/spider_collection · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2402
  • days_rel: n/a
  • days_push: 863
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1646 stars · 239 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A collection of Python web crawler scripts targeting sites like Bilibili, Zhihu, Weibo, NetEase Music, GitHub, and Anjuke, built with requests, Scrapy, parsel, and BeautifulSoup. It also includes a re-wrapped IP proxy pool and simple data analysis/visualization examples using pandas and matplotlib.

Use cases

  • download bilibili videos or extract video covers
  • scrape zhihu answers and user profiles for analysis
  • collect weibo user info
  • download netease cloud music playlists
  • scrape rental listings from anjuke and visualize them
  • learn how to build python crawlers with scrapy and requests
  • set up a wrapped ip proxy pool for anti-scraping

When to choose

  • you want ready-made example crawlers for popular Chinese websites
  • you are learning web scraping techniques like threading, proxy pools, and scrapy
  • you need a starting point to adapt a crawler for bilibili, zhihu, or weibo

When to avoid

  • you need a production-grade, maintained scraping framework
  • you need a general-purpose crawler with a stable API
  • the target site has changed its layout and the script is outdated (some spiders are marked deprecated)

Facets

library · maturity active

web-scraping data-visualization caching developer-tools crawlers data-science social-media developer-tools python cross-platform scrapy requests proxy-pool bilibili zhihu weibo netease-music multi-spider-collection learning-project automation

1 source

Member repositories

RepositoryRoleHealth v2
srx-2000/spider_collectionmain32

For agents

markdown · JSON · MCP: product_card(name="srx-2000/spider_collection")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem