Ross ROSS = Recommend OSS · open-source software intelligence for agents

resource: web-scraping

198 resources, primary matches first, then adoption-weighted; health v2 shown.

ResourceHealth v2StarsMaturity
JCodesMore/ai-website-cloner-template
A template repository that lets AI coding agents like Claude Code recreate any website from a URL as a clean Next.js app with one command. …
8033188active
StevenBlack/hosts
A hosts file aggregator that consolidates and merges several well-curated hosts files into unified, deduplicated hosts files for DNS-level …
9530954active
fanmingming/live
A free, openly hosted library of TV and radio channel logos (icons) plus related IPTV tools including EPG XML, M3U/TXT playlist converters,…
7728387active
timqian/chinese-independent-blogs
A curated list of Chinese independent blogs with RSS feeds, ranked by RSS subscription counts. It serves as a discovery directory for Chine…
7723859active
wistbean/learn_python3_spider
A Chinese-language tutorial series and example code repository teaching Python web scraping from zero to advanced, covering packet capture …
7522033active
jbiaojerry/ebook-treasure-chest
A curated collection of ebook download links (epub, mobi, azw3) organized by category, sourced from Chinese reading apps like Fan Deng Read…
5616575active
Chinese DOS Games
A curated collection of 1,898 Chinese-language DOS games that can be downloaded via a Python script and played in the browser through an Em…
3210291active
joevess/IPTV
An automatically updated collection of IPTV live-stream playlists (m3u8) aggregating sources from haoqu, TVBox, and other public sources, s…
2810276active
hoochanlon/hamuleite
A curated knowledge-base repository aggregating links, articles, and academic paper resources in social sciences, economics, mathematics, g…
779566active
jackvale/rectg
A curated index of 500+ Chinese-language Telegram channels, groups, and bots, organized into 22 topic categories with automated scraping pl…
709182active
cipher387/osint_stuff_tool_collection
A curated awesome-list collection of 1000+ online tools for OSINT (open-source intelligence), organized into categories like geolocation, s…
698738active
lorien/awesome-web-scraping
A curated awesome-list of web scraping libraries, tools, APIs, and manuals across Python, PHP, Ruby, JavaScript, Go, and CLI tools. It also…
778133active
Rockyzsu/stock
A continuously updated Chinese-language tutorial series and code collection for learning quantitative stock trading in 30 days, built aroun…
677965active
snehasishroy/leetcode-companywise-interview-questions
A curated dataset of LeetCode questions categorized by company (Google, Amazon, Meta, Microsoft, etc.) and recency, with difficulty, accept…
767549active
cporter202/API-mega-list
A large curated directory of over 11,000 public APIs organized into 24 categories, maintained as a GitHub awesome-list. It serves as a refe…
587532active
xiangyuecn/AreaCity-JsSpider-StatsGov
A dataset and tooling project providing China's province/city/district/town (3-4 level) administrative division data with pinyin, coordinat…
706842active
dhamaniasad/HeadlessBrowsers
A curated list of almost all headless web browsers in existence, covering browser engines, multi-driver libraries, and language-specific to…
536683active
proxifly/free-proxy-list
A continuously updated free proxy list (HTTP, HTTPS, SOCKS4, SOCKS5) refreshed every 5 minutes, sourced from 100+ countries and available i…
736643active
reddelexc/hackerone-reports
A curated dataset of top disclosed HackerOne bug bounty reports, ranked by upvotes, bounties, bug type, and program, with raw data in data.…
766481active
aoaostar/legado
A curated collection of book sources, subscription feeds, themes, and layout configs for the Legado (阅读) Android reading app, served via a …
766098active
grapeot/devin.cursorrules
A configuration template and toolset that turns Cursor, Windsurf, or GitHub Copilot into a Devin-like agentic AI coding assistant via .curs…
315970active
anbeime/skill
A curated AI agent skill store aggregating hundreds of packaged skills (for Claude, Gemini, and other agents) covering document processing,…
605815active
TheSpeedX/PROXY-List
A regularly updated dataset of free public proxy servers, provided as plain-text lists of SOCKS4, SOCKS5, and HTTP proxies. It aggregates t…
775780active
niespodd/browser-fingerprinting
An educational guide and analysis repository explaining how bot protection systems (PerimeterX, Akamai, Kasada, Arkose, reCAPTCHA, etc.) de…
755125active
gayanvoice/top-github-users
A continuously updated dataset and leaderboard of the most active GitHub users ranked by public/private contributions and followers, organi…
774884active
hiddendevj/Crawler_Illegal_Cases_In_China
A curated collection of legal cases, news, and Chinese laws related to web crawler developers facing prosecution or violations in mainland …
644717active
WeNeedHome/SummaryOfLoanSuspension
A crowdsourced, open dataset aggregating mortgage suspension (loan boycott) notices from unfinished housing projects across Chinese provinc…
3220366maintenance
Doragd/Algorithm-Practice-in-Industry
A curated collection of industry practice articles, top-conference papers, and blog posts on search, recommendation, and advertising (搜广推) …
754574active
NanmiCoder/CrawlerTutorial
A Chinese-language open-source tutorial series teaching web scraping from beginner to advanced levels, written by the author of MediaCrawle…
594552active
wpzzz/blocked-sites-in-south-korea
A dataset tracking websites blocked by the South Korean government, maintained via Python scripts. It provides a regularly updated list of …
324541active
Jack-Cherish/python-spider
A collection of Python3 web scraping example scripts and tutorials covering sites like Taobao, JD, Bilibili, 12306, Douyin, and novel/comic…
3219741maintenance
shashankvemuri/Finance
A collection of 150+ standalone Python programs for gathering, manipulating, and analyzing stock market data. It covers stock screening, ma…
664193active
ai-robots-txt/ai.robots.txt
A community-maintained list of AI crawler user agents, distributed as robots.txt plus ready-made blocking configs for Apache, Nginx, Caddy,…
924082active
cporter202/scraping-apis-for-devs
A curated directory of 2,622 scraping APIs organized into 17 categories, covering data extraction from websites, social media, and e-commer…
443847active
xiaohucode/yidaRule
A community rule repository ('Yida Rules') for the Yida app, a cross-platform Flutter + Rust media aggregator that plays video, audio, read…
103811active
Kr1s77/awesome-python-login-model
A collection of Python example scripts demonstrating how to simulate logins on major Chinese and international websites using Selenium, Web…
3216218maintenance
avinashkranjan/Amazing-Python-Scripts
A curated collection of Python scripts ranging from basic to advanced, including automation task scripts, contributed by the open-source co…
633662active
w3c/IntersectionObserver
The W3C specification repository for the Intersection Observer browser API, written in Bikeshed, which defines how web pages efficiently ob…
653613active
Yixiaohan/show-me-the-code
A curated collection of small daily Python programming exercises designed for learners to practice coding skills. Each exercise is a self-c…
3213729maintenance
dwyl/learn-to-send-email-via-google-script-html-no-server
A step-by-step tutorial showing how to send email from a static HTML form (e.g. a 'Contact Us' page) using Google Apps Script, with no back…
323210active
oxylabs/how-to-scrape-amazon-product-data
A Python tutorial repository from Oxylabs demonstrating how to scrape Amazon product data (titles, ratings, prices, images, descriptions) u…
613197active
alex000kim/nsfw_data_scraper
A collection of shell scripts that automatically aggregate tens of thousands of images across five categories (porn, hentai, sexy, neutral,…
3212588maintenance
x4nth055/pythoncode-tutorials
A large collection of Jupyter Notebook and Python code examples accompanying the tutorials from ThePythonCode.com website. It covers topics…
743000active
XIU2/Yuedu
A curated collection of book source rules for the Legado (阅读) Android e-reader app, which parses third-party novel websites for search, det…
7612135maintenance
Jack-Cherish/PythonPark
A curated Chinese-language collection of Python self-study tutorials covering machine learning, deep learning, web scraping, data structure…
3211677maintenance
oxylabs/how-to-scrape-google-trends
A step-by-step tutorial repository showing how to scrape Google Trends data (keywords, popularity, regional breakdown, related queries) usi…
502853active
hi-weijun/PythonDataScience-Collections
A curated Chinese-language collection of links and resources for Python data analysis, covering Python basics, web scraping, visualization,…
692795active
pibigstar/go-demo
A Go language example tutorial repository covering basics through advanced topics, including standard library usage, design patterns, inter…
342709active
injetlee/Python
A collection of Python example scripts and tutorial-style code covering web scraping, simulated logins (e.g., Zhihu), Excel file reading/wr…
6910800maintenance
iipc/awesome-web-archiving
A curated Awesome List of resources for getting started with web archiving, covering training materials, the WARC standard, and tools for a…
762626active
apurvsinghgautam/dark-web-osint-tools
A curated list of open-source OSINT tools for the dark web, organized by category: search engines, onion link discovery, onion link scannin…
752550active
cipher387/API-s-for-OSINT
A curated awesome-list of APIs useful for automating OSINT (open-source intelligence) tasks, covering phone number lookup, domain/DNS/IP lo…
752507active
larymak/Python-project-Scripts
A curated collection of beginner-level Python script projects maintained as an open-source learning repository. Contributors add small stan…
762472active
puppeteer/examples
A collection of use case-driven JavaScript examples demonstrating how to use Puppeteer and headless Chrome for browser automation tasks. Ea…
722415active
clarketm/proxy-list
A daily-updated list of free, public forward proxy servers with metadata on country, anonymity level, protocol type, and Google pass status…
642385active
zhaoolee/ins
A curated, ad-free database of inspiring websites for internet professionals, stored as CSV and rendered as a browsable catalog. GitHub Act…
772254active
zilong7728/Collect-IPTV
An automated IPTV channel source collection project that aggregates publicly available live TV streams, tests their availability and latenc…
652222active
talkpython/100daysofcode-with-python-course
Course materials and handouts for the Talk Python #100DaysOfCode in Python course, containing 33 guided projects across 100 days of learnin…
322203stable
xianhu/LearnPython
A collection of Python learning scripts and Jupyter notebooks that teach Python concepts through runnable code examples, from basics to adv…
328560maintenance
cporter202/social-media-scraping-apis
A curated catalog of thousands of third-party social media scraping APIs for extracting posts, profiles, videos, comments, and engagement m…
442198active
Jieyab89/OSINT-Cheat-sheet
A curated cheat sheet repository listing OSINT (open-source intelligence) tools, datasets, wikis, articles, and tutorials for reconnaissanc…
772182active
BlueSkyXN/AdGuardHomeRules
A large aggregated collection of AdGuard Home DNS ad-blocking and filtering rules, with blacklist and whitelist lists maintained via automa…
492134active
tinyfish-io/tinyfish-cookbook
A collection of open-source sample apps, recipes, and demos built on the TinyFish web agent platform, which provides Search, Fetch, Agent, …
602120active
NateScarlet/holiday-cn
A machine-readable dataset of China's official statutory holidays, automatically scraped daily from State Council announcements and publish…
782105active
scraly/developers-conferences-agenda
A community-driven, open dataset and web platform listing developer/tech conferences and Calls for Papers (CFPs) worldwide, viewable as a l…
672000active
lixi5338619/lxSpider
A collection of Python web scraping example scripts covering many Chinese platforms (Taobao, Douyin, Weibo, WeChat, Xiaohongshu, etc.) plus…
231968active
oxylabs/how-to-scrape-google-scholar
A tutorial repository with example Python code showing how to scrape Google Scholar results (titles, authors, citation counts, PDF links) u…
691963active
mursor1985/LIVE
A community-maintained collection of live TV / IPTV stream sources (m3u-style playlists) aggregated from the internet, primarily Chinese-la…
621961active
Momo707577045/media-source-extract
A tutorial and tool for indiscriminately extracting videos that use MediaSource (MSE) playback, by intercepting video segments at the final…
401955active
BruceDone/awesome-crawler
A curated awesome-list cataloging web crawler, spider, and scraper tools across many programming languages including Python, Java, JavaScri…
327296maintenance
luyishisi/Anti-Anti-Spider
A Chinese-language repository collecting techniques and code for bypassing anti-scraping measures on websites, including a CNN-based (AlexN…
327277maintenance
hyperbrowserai/hyperbrowser-app-examples
A collection of complete, production-ready example web applications built with Hyperbrowser, a browser automation and web scraping platform…
621909active
TheGP/untidetect-tools
A curated list of anti-detect browsers, humanizing tools, captcha solvers, and SMS activation services, maintained as a reference for brows…
681899active
u3c3/BT-btt
A documentation repository that tracks the latest mirror domains and usage guidance for the U3C3 magnet link website, an adult-focused BitT…
231885active
LawRefBook/Laws
A curated dataset of Chinese laws, regulations, and departmental rules, structured by chapters and stored in Markdown with a generated SQLi…
631845active
yhangf/PythonCrawler
A collection of Python web crawler scripts covering tasks like scraping images from Baidu, job postings, JD product data, and GitHub trendi…
691820active
trickest/wordlists
A regularly updated collection of real-world infosec wordlists maintained by Trickest, including technology-specific path lists (WordPress,…
771791active
openai/openai-cua-sample-app
A TypeScript sample application from OpenAI demonstrating how to use the Computer Using Agent (CUA) via the Responses API against browser e…
531775active
xishandong/crawlProject
A collection of hands-on Python web scraping practice projects ranging from beginner requests-based crawlers to JavaScript reverse engineer…
281769active
TheWebScrapingClub/webscraping-from-0-to-hero
A community-driven knowledge repository about web scraping with Python, curated by The Web Scraping Club newsletter author. It aggregates g…
321735active
icopy-site/awesome-cn
A curated aggregation of GitHub awesome lists, collected by a scheduled crawler and published as a Chinese-language documentation site buil…
751727active
oxylabs/how-to-scrape-google-jobs
A tutorial repository with sample Python code showing how to scrape Google Jobs listings, both with a free scraper and at scale using Oxyla…
651707active
oxylabs/how-to-handle-amazon-captcha
A tutorial repository with Python example code showing how to handle CAPTCHAs when scraping Amazon product data, comparing a plain requests…
651701active
RPiList/specials
A curated collection of DNS blocklists for Pi-hole that protect against fake shops, advertising, tracking, and other internet threats. It a…
771692active
oxylabs/scrape-google-python
A Python tutorial repository from Oxylabs demonstrating how to scrape Google search results (SERPs) using Oxylabs' SERP Scraper API. It inc…
631629active
tanjiti/sec_profile
A continuously updated dataset and scraping project that crawls security information sources like secwiki and xuanwu.github.io/sec.today, a…
771602active
mxschmitt/awesome-playwright
A curated awesome-list of tools, utilities, integrations, and projects built around the Playwright browser automation and testing framework…
761560active
DropsDevopsOrg/ECommerceCrawlers
A curated collection of Python web crawler projects targeting Chinese e-commerce and content websites such as Taobao, Dianping, WeChat publ…
235664maintenance
zapplyjobs/New-Grad-Jobs-2027
A continuously updated, community-curated job board listing new grad, intern, and early-career roles across tech, finance, healthcare, and …
771519active
Shrans/GalSites
A curated directory of free Galgame (visual novel) resource sites, maintained as a README list with community contributions via Issues and …
661507active
monosans/proxy-list
A continuously updated dataset of free HTTP, SOCKS4, and SOCKS5 proxies, re-verified every hour and published as plain text and JSON files …
771498active
gxcuizy/Python
A collection of Python 3 example programs and small scripts, including a learn-Python-from-scratch series, a 12306 train ticket grabbing sc…
325398maintenance
ArthurHeitmann/arctic_shift
Project Arctic Shift is an archive of Reddit data (posts, comments) made accessible through large compressed dumps, a limited API, and a we…
931438active
darbra/sperm
A curated collection of reverse-engineering articles gathered from Chinese platforms like 52pojie, Kanxue, CSDN, and WeChat public accounts…
671409active
carlospolop/Auto_Wordlists
A repository of automatically generated security wordlists for web fuzzing, reconnaissance, and payload testing, refreshed on a schedule (D…
761403active
Alfred1984/interesting-python
A collection of small, fun Python projects demonstrating web scraping and data analysis, written as Jupyter Notebooks and paired with Chine…
325007maintenance
entr0pia/SwitchyOmega-Whitelist
An auto-updated whitelist of Chinese mainland domains formatted for the SwitchyOmega/ZeroOmega browser proxy extension, sourced from dnsmas…
671387active
kgspider/crawler
A collection of example code accompanying the 'K哥爬虫' tutorial series on JavaScript reverse engineering for web scraping. Each subdirectory …
421379active
oxylabs/how-to-scrape-google-finance
A Python tutorial repository from Oxylabs demonstrating how to scrape Google Finance data (stock titles, prices, and percentage price chang…
421346active
REMitchell/python-scraping
Companion code samples for the O'Reilly book 'Web Scraping with Python' (2nd Edition), mostly provided as Jupyter notebooks. It covers scra…
324723maintenance

page 1 / 2 next →