resource: web-scraping
198 resources, primary matches first, then adoption-weighted; health v2 shown.
| Resource | Health v2 | Stars | Maturity |
|---|---|---|---|
| JCodesMore/ai-website-cloner-template A template repository that lets AI coding agents like Claude Code recreate any website from a URL as a clean Next.js app with one command. … | 80 | 33188 | active |
| StevenBlack/hosts A hosts file aggregator that consolidates and merges several well-curated hosts files into unified, deduplicated hosts files for DNS-level … | 95 | 30954 | active |
| fanmingming/live A free, openly hosted library of TV and radio channel logos (icons) plus related IPTV tools including EPG XML, M3U/TXT playlist converters,… | 77 | 28387 | active |
| timqian/chinese-independent-blogs A curated list of Chinese independent blogs with RSS feeds, ranked by RSS subscription counts. It serves as a discovery directory for Chine… | 77 | 23859 | active |
| wistbean/learn_python3_spider A Chinese-language tutorial series and example code repository teaching Python web scraping from zero to advanced, covering packet capture … | 75 | 22033 | active |
| jbiaojerry/ebook-treasure-chest A curated collection of ebook download links (epub, mobi, azw3) organized by category, sourced from Chinese reading apps like Fan Deng Read… | 56 | 16575 | active |
| Chinese DOS Games A curated collection of 1,898 Chinese-language DOS games that can be downloaded via a Python script and played in the browser through an Em… | 32 | 10291 | active |
| joevess/IPTV An automatically updated collection of IPTV live-stream playlists (m3u8) aggregating sources from haoqu, TVBox, and other public sources, s… | 28 | 10276 | active |
| hoochanlon/hamuleite A curated knowledge-base repository aggregating links, articles, and academic paper resources in social sciences, economics, mathematics, g… | 77 | 9566 | active |
| jackvale/rectg A curated index of 500+ Chinese-language Telegram channels, groups, and bots, organized into 22 topic categories with automated scraping pl… | 70 | 9182 | active |
| cipher387/osint_stuff_tool_collection A curated awesome-list collection of 1000+ online tools for OSINT (open-source intelligence), organized into categories like geolocation, s… | 69 | 8738 | active |
| lorien/awesome-web-scraping A curated awesome-list of web scraping libraries, tools, APIs, and manuals across Python, PHP, Ruby, JavaScript, Go, and CLI tools. It also… | 77 | 8133 | active |
| Rockyzsu/stock A continuously updated Chinese-language tutorial series and code collection for learning quantitative stock trading in 30 days, built aroun… | 67 | 7965 | active |
| snehasishroy/leetcode-companywise-interview-questions A curated dataset of LeetCode questions categorized by company (Google, Amazon, Meta, Microsoft, etc.) and recency, with difficulty, accept… | 76 | 7549 | active |
| cporter202/API-mega-list A large curated directory of over 11,000 public APIs organized into 24 categories, maintained as a GitHub awesome-list. It serves as a refe… | 58 | 7532 | active |
| xiangyuecn/AreaCity-JsSpider-StatsGov A dataset and tooling project providing China's province/city/district/town (3-4 level) administrative division data with pinyin, coordinat… | 70 | 6842 | active |
| dhamaniasad/HeadlessBrowsers A curated list of almost all headless web browsers in existence, covering browser engines, multi-driver libraries, and language-specific to… | 53 | 6683 | active |
| proxifly/free-proxy-list A continuously updated free proxy list (HTTP, HTTPS, SOCKS4, SOCKS5) refreshed every 5 minutes, sourced from 100+ countries and available i… | 73 | 6643 | active |
| reddelexc/hackerone-reports A curated dataset of top disclosed HackerOne bug bounty reports, ranked by upvotes, bounties, bug type, and program, with raw data in data.… | 76 | 6481 | active |
| aoaostar/legado A curated collection of book sources, subscription feeds, themes, and layout configs for the Legado (阅读) Android reading app, served via a … | 76 | 6098 | active |
| grapeot/devin.cursorrules A configuration template and toolset that turns Cursor, Windsurf, or GitHub Copilot into a Devin-like agentic AI coding assistant via .curs… | 31 | 5970 | active |
| anbeime/skill A curated AI agent skill store aggregating hundreds of packaged skills (for Claude, Gemini, and other agents) covering document processing,… | 60 | 5815 | active |
| TheSpeedX/PROXY-List A regularly updated dataset of free public proxy servers, provided as plain-text lists of SOCKS4, SOCKS5, and HTTP proxies. It aggregates t… | 77 | 5780 | active |
| niespodd/browser-fingerprinting An educational guide and analysis repository explaining how bot protection systems (PerimeterX, Akamai, Kasada, Arkose, reCAPTCHA, etc.) de… | 75 | 5125 | active |
| gayanvoice/top-github-users A continuously updated dataset and leaderboard of the most active GitHub users ranked by public/private contributions and followers, organi… | 77 | 4884 | active |
| hiddendevj/Crawler_Illegal_Cases_In_China A curated collection of legal cases, news, and Chinese laws related to web crawler developers facing prosecution or violations in mainland … | 64 | 4717 | active |
| WeNeedHome/SummaryOfLoanSuspension A crowdsourced, open dataset aggregating mortgage suspension (loan boycott) notices from unfinished housing projects across Chinese provinc… | 32 | 20366 | maintenance |
| Doragd/Algorithm-Practice-in-Industry A curated collection of industry practice articles, top-conference papers, and blog posts on search, recommendation, and advertising (搜广推) … | 75 | 4574 | active |
| NanmiCoder/CrawlerTutorial A Chinese-language open-source tutorial series teaching web scraping from beginner to advanced levels, written by the author of MediaCrawle… | 59 | 4552 | active |
| wpzzz/blocked-sites-in-south-korea A dataset tracking websites blocked by the South Korean government, maintained via Python scripts. It provides a regularly updated list of … | 32 | 4541 | active |
| Jack-Cherish/python-spider A collection of Python3 web scraping example scripts and tutorials covering sites like Taobao, JD, Bilibili, 12306, Douyin, and novel/comic… | 32 | 19741 | maintenance |
| shashankvemuri/Finance A collection of 150+ standalone Python programs for gathering, manipulating, and analyzing stock market data. It covers stock screening, ma… | 66 | 4193 | active |
| ai-robots-txt/ai.robots.txt A community-maintained list of AI crawler user agents, distributed as robots.txt plus ready-made blocking configs for Apache, Nginx, Caddy,… | 92 | 4082 | active |
| cporter202/scraping-apis-for-devs A curated directory of 2,622 scraping APIs organized into 17 categories, covering data extraction from websites, social media, and e-commer… | 44 | 3847 | active |
| xiaohucode/yidaRule A community rule repository ('Yida Rules') for the Yida app, a cross-platform Flutter + Rust media aggregator that plays video, audio, read… | 10 | 3811 | active |
| Kr1s77/awesome-python-login-model A collection of Python example scripts demonstrating how to simulate logins on major Chinese and international websites using Selenium, Web… | 32 | 16218 | maintenance |
| avinashkranjan/Amazing-Python-Scripts A curated collection of Python scripts ranging from basic to advanced, including automation task scripts, contributed by the open-source co… | 63 | 3662 | active |
| w3c/IntersectionObserver The W3C specification repository for the Intersection Observer browser API, written in Bikeshed, which defines how web pages efficiently ob… | 65 | 3613 | active |
| Yixiaohan/show-me-the-code A curated collection of small daily Python programming exercises designed for learners to practice coding skills. Each exercise is a self-c… | 32 | 13729 | maintenance |
| dwyl/learn-to-send-email-via-google-script-html-no-server A step-by-step tutorial showing how to send email from a static HTML form (e.g. a 'Contact Us' page) using Google Apps Script, with no back… | 32 | 3210 | active |
| oxylabs/how-to-scrape-amazon-product-data A Python tutorial repository from Oxylabs demonstrating how to scrape Amazon product data (titles, ratings, prices, images, descriptions) u… | 61 | 3197 | active |
| alex000kim/nsfw_data_scraper A collection of shell scripts that automatically aggregate tens of thousands of images across five categories (porn, hentai, sexy, neutral,… | 32 | 12588 | maintenance |
| x4nth055/pythoncode-tutorials A large collection of Jupyter Notebook and Python code examples accompanying the tutorials from ThePythonCode.com website. It covers topics… | 74 | 3000 | active |
| XIU2/Yuedu A curated collection of book source rules for the Legado (阅读) Android e-reader app, which parses third-party novel websites for search, det… | 76 | 12135 | maintenance |
| Jack-Cherish/PythonPark A curated Chinese-language collection of Python self-study tutorials covering machine learning, deep learning, web scraping, data structure… | 32 | 11677 | maintenance |
| oxylabs/how-to-scrape-google-trends A step-by-step tutorial repository showing how to scrape Google Trends data (keywords, popularity, regional breakdown, related queries) usi… | 50 | 2853 | active |
| hi-weijun/PythonDataScience-Collections A curated Chinese-language collection of links and resources for Python data analysis, covering Python basics, web scraping, visualization,… | 69 | 2795 | active |
| pibigstar/go-demo A Go language example tutorial repository covering basics through advanced topics, including standard library usage, design patterns, inter… | 34 | 2709 | active |
| injetlee/Python A collection of Python example scripts and tutorial-style code covering web scraping, simulated logins (e.g., Zhihu), Excel file reading/wr… | 69 | 10800 | maintenance |
| iipc/awesome-web-archiving A curated Awesome List of resources for getting started with web archiving, covering training materials, the WARC standard, and tools for a… | 76 | 2626 | active |
| apurvsinghgautam/dark-web-osint-tools A curated list of open-source OSINT tools for the dark web, organized by category: search engines, onion link discovery, onion link scannin… | 75 | 2550 | active |
| cipher387/API-s-for-OSINT A curated awesome-list of APIs useful for automating OSINT (open-source intelligence) tasks, covering phone number lookup, domain/DNS/IP lo… | 75 | 2507 | active |
| larymak/Python-project-Scripts A curated collection of beginner-level Python script projects maintained as an open-source learning repository. Contributors add small stan… | 76 | 2472 | active |
| puppeteer/examples A collection of use case-driven JavaScript examples demonstrating how to use Puppeteer and headless Chrome for browser automation tasks. Ea… | 72 | 2415 | active |
| clarketm/proxy-list A daily-updated list of free, public forward proxy servers with metadata on country, anonymity level, protocol type, and Google pass status… | 64 | 2385 | active |
| zhaoolee/ins A curated, ad-free database of inspiring websites for internet professionals, stored as CSV and rendered as a browsable catalog. GitHub Act… | 77 | 2254 | active |
| zilong7728/Collect-IPTV An automated IPTV channel source collection project that aggregates publicly available live TV streams, tests their availability and latenc… | 65 | 2222 | active |
| talkpython/100daysofcode-with-python-course Course materials and handouts for the Talk Python #100DaysOfCode in Python course, containing 33 guided projects across 100 days of learnin… | 32 | 2203 | stable |
| xianhu/LearnPython A collection of Python learning scripts and Jupyter notebooks that teach Python concepts through runnable code examples, from basics to adv… | 32 | 8560 | maintenance |
| cporter202/social-media-scraping-apis A curated catalog of thousands of third-party social media scraping APIs for extracting posts, profiles, videos, comments, and engagement m… | 44 | 2198 | active |
| Jieyab89/OSINT-Cheat-sheet A curated cheat sheet repository listing OSINT (open-source intelligence) tools, datasets, wikis, articles, and tutorials for reconnaissanc… | 77 | 2182 | active |
| BlueSkyXN/AdGuardHomeRules A large aggregated collection of AdGuard Home DNS ad-blocking and filtering rules, with blacklist and whitelist lists maintained via automa… | 49 | 2134 | active |
| tinyfish-io/tinyfish-cookbook A collection of open-source sample apps, recipes, and demos built on the TinyFish web agent platform, which provides Search, Fetch, Agent, … | 60 | 2120 | active |
| NateScarlet/holiday-cn A machine-readable dataset of China's official statutory holidays, automatically scraped daily from State Council announcements and publish… | 78 | 2105 | active |
| scraly/developers-conferences-agenda A community-driven, open dataset and web platform listing developer/tech conferences and Calls for Papers (CFPs) worldwide, viewable as a l… | 67 | 2000 | active |
| lixi5338619/lxSpider A collection of Python web scraping example scripts covering many Chinese platforms (Taobao, Douyin, Weibo, WeChat, Xiaohongshu, etc.) plus… | 23 | 1968 | active |
| oxylabs/how-to-scrape-google-scholar A tutorial repository with example Python code showing how to scrape Google Scholar results (titles, authors, citation counts, PDF links) u… | 69 | 1963 | active |
| mursor1985/LIVE A community-maintained collection of live TV / IPTV stream sources (m3u-style playlists) aggregated from the internet, primarily Chinese-la… | 62 | 1961 | active |
| Momo707577045/media-source-extract A tutorial and tool for indiscriminately extracting videos that use MediaSource (MSE) playback, by intercepting video segments at the final… | 40 | 1955 | active |
| BruceDone/awesome-crawler A curated awesome-list cataloging web crawler, spider, and scraper tools across many programming languages including Python, Java, JavaScri… | 32 | 7296 | maintenance |
| luyishisi/Anti-Anti-Spider A Chinese-language repository collecting techniques and code for bypassing anti-scraping measures on websites, including a CNN-based (AlexN… | 32 | 7277 | maintenance |
| hyperbrowserai/hyperbrowser-app-examples A collection of complete, production-ready example web applications built with Hyperbrowser, a browser automation and web scraping platform… | 62 | 1909 | active |
| TheGP/untidetect-tools A curated list of anti-detect browsers, humanizing tools, captcha solvers, and SMS activation services, maintained as a reference for brows… | 68 | 1899 | active |
| u3c3/BT-btt A documentation repository that tracks the latest mirror domains and usage guidance for the U3C3 magnet link website, an adult-focused BitT… | 23 | 1885 | active |
| LawRefBook/Laws A curated dataset of Chinese laws, regulations, and departmental rules, structured by chapters and stored in Markdown with a generated SQLi… | 63 | 1845 | active |
| yhangf/PythonCrawler A collection of Python web crawler scripts covering tasks like scraping images from Baidu, job postings, JD product data, and GitHub trendi… | 69 | 1820 | active |
| trickest/wordlists A regularly updated collection of real-world infosec wordlists maintained by Trickest, including technology-specific path lists (WordPress,… | 77 | 1791 | active |
| openai/openai-cua-sample-app A TypeScript sample application from OpenAI demonstrating how to use the Computer Using Agent (CUA) via the Responses API against browser e… | 53 | 1775 | active |
| xishandong/crawlProject A collection of hands-on Python web scraping practice projects ranging from beginner requests-based crawlers to JavaScript reverse engineer… | 28 | 1769 | active |
| TheWebScrapingClub/webscraping-from-0-to-hero A community-driven knowledge repository about web scraping with Python, curated by The Web Scraping Club newsletter author. It aggregates g… | 32 | 1735 | active |
| icopy-site/awesome-cn A curated aggregation of GitHub awesome lists, collected by a scheduled crawler and published as a Chinese-language documentation site buil… | 75 | 1727 | active |
| oxylabs/how-to-scrape-google-jobs A tutorial repository with sample Python code showing how to scrape Google Jobs listings, both with a free scraper and at scale using Oxyla… | 65 | 1707 | active |
| oxylabs/how-to-handle-amazon-captcha A tutorial repository with Python example code showing how to handle CAPTCHAs when scraping Amazon product data, comparing a plain requests… | 65 | 1701 | active |
| RPiList/specials A curated collection of DNS blocklists for Pi-hole that protect against fake shops, advertising, tracking, and other internet threats. It a… | 77 | 1692 | active |
| oxylabs/scrape-google-python A Python tutorial repository from Oxylabs demonstrating how to scrape Google search results (SERPs) using Oxylabs' SERP Scraper API. It inc… | 63 | 1629 | active |
| tanjiti/sec_profile A continuously updated dataset and scraping project that crawls security information sources like secwiki and xuanwu.github.io/sec.today, a… | 77 | 1602 | active |
| mxschmitt/awesome-playwright A curated awesome-list of tools, utilities, integrations, and projects built around the Playwright browser automation and testing framework… | 76 | 1560 | active |
| DropsDevopsOrg/ECommerceCrawlers A curated collection of Python web crawler projects targeting Chinese e-commerce and content websites such as Taobao, Dianping, WeChat publ… | 23 | 5664 | maintenance |
| zapplyjobs/New-Grad-Jobs-2027 A continuously updated, community-curated job board listing new grad, intern, and early-career roles across tech, finance, healthcare, and … | 77 | 1519 | active |
| Shrans/GalSites A curated directory of free Galgame (visual novel) resource sites, maintained as a README list with community contributions via Issues and … | 66 | 1507 | active |
| monosans/proxy-list A continuously updated dataset of free HTTP, SOCKS4, and SOCKS5 proxies, re-verified every hour and published as plain text and JSON files … | 77 | 1498 | active |
| gxcuizy/Python A collection of Python 3 example programs and small scripts, including a learn-Python-from-scratch series, a 12306 train ticket grabbing sc… | 32 | 5398 | maintenance |
| ArthurHeitmann/arctic_shift Project Arctic Shift is an archive of Reddit data (posts, comments) made accessible through large compressed dumps, a limited API, and a we… | 93 | 1438 | active |
| darbra/sperm A curated collection of reverse-engineering articles gathered from Chinese platforms like 52pojie, Kanxue, CSDN, and WeChat public accounts… | 67 | 1409 | active |
| carlospolop/Auto_Wordlists A repository of automatically generated security wordlists for web fuzzing, reconnaissance, and payload testing, refreshed on a schedule (D… | 76 | 1403 | active |
| Alfred1984/interesting-python A collection of small, fun Python projects demonstrating web scraping and data analysis, written as Jupyter Notebooks and paired with Chine… | 32 | 5007 | maintenance |
| entr0pia/SwitchyOmega-Whitelist An auto-updated whitelist of Chinese mainland domains formatted for the SwitchyOmega/ZeroOmega browser proxy extension, sourced from dnsmas… | 67 | 1387 | active |
| kgspider/crawler A collection of example code accompanying the 'K哥爬虫' tutorial series on JavaScript reverse engineering for web scraping. Each subdirectory … | 42 | 1379 | active |
| oxylabs/how-to-scrape-google-finance A Python tutorial repository from Oxylabs demonstrating how to scrape Google Finance data (stock titles, prices, and percentage price chang… | 42 | 1346 | active |
| REMitchell/python-scraping Companion code samples for the O'Reilly book 'Web Scraping with Python' (2nd Edition), mostly provided as Jupyter notebooks. It covers scra… | 32 | 4723 | maintenance |
page 1 / 2 next →