function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| jmcarp/robobrowser RoboBrowser is a Pythonic library for browsing the web without a standalone browser, combining Requests for HTTP sessions with BeautifulSou… | 32 | 3692 | maintenance |
| jae-jae/fetcher-mcp A Model Context Protocol (MCP) server that fetches web page content using a Playwright headless browser, executing JavaScript to handle dyn… | 48 | 1076 | active |
| linkchecker/linkchecker LinkChecker is a GPL-licensed Python tool that checks links in web documents or entire websites for broken URLs. It supports recursive mult… | 65 | 1075 | active |
| 1061700625/WeChat_Article A PyQt5 desktop application that crawls and downloads all articles from a specified WeChat official account. It uses Selenium to log in and… | 62 | 1075 | active |
| scrapfly/scrapfly-scrapers A collection of educational Python web scraping scripts for over 40 popular domains such as Amazon, AliExpress, BestBuy, and Twitter, built… | 74 | 1074 | active |
| xuejianxianzun/PixivFanboxDownloader A Chrome browser extension for batch downloading files from Pixiv Fanbox posts. It supports file type filtering, custom filenames, multiple… | 98 | 1073 | active |
| soxoj/socid-extractor socid_extractor is a Python library and CLI that extracts structured account metadata and stable internal identifiers (usernames, UIDs, GAI… | 92 | 1073 | active |
| Junyi-99/ChatGPT-API-Scanner A Python CLI tool that scans GitHub for publicly leaked OpenAI API keys using Selenium browser automation. It is intended for security rese… | 58 | 1073 | active |
| tomnomnom/assetfinder A Go command-line tool that discovers domains and subdomains potentially related to a given domain by querying multiple passive sources lik… | 23 | 3666 | maintenance |
| borisbabic/browser_cookie3 A Python library (fork of browsercookie) that loads cookies from installed web browsers like Chrome, Firefox, Edge, Safari, and others into… | 23 | 1070 | active |
| jikan-me/jikan Jikan is an unofficial PHP library and REST API for MyAnimeList.net that scrapes the website to provide data the official API lacks. It let… | 64 | 1066 | active |
| QingJ01/123pan_unlock A Tampermonkey userscript that unlocks download restrictions on the 123pan (123云盘) cloud storage service, including bypassing the 1GB downl… | 10 | 1062 | active |
| jez500/pricebuddy PriceBuddy is a self-hostable web application that tracks product prices and availability across online stores, keeping price history and n… | 86 | 1060 | active |
| unitedstates/congress A community-run Python toolkit that collects and converts official U.S. Congress data—bills, amendments, roll call votes, nominations, and … | 52 | 1060 | active |
| cantino/selectorgadget SelectorGadget is an open-source bookmarklet and Chrome extension that generates CSS selectors for page elements through point-and-click se… | 69 | 1059 | active |
| spider-ios/autox-release AutoX is a desktop social media operations tool that automates one-click publishing of videos to multiple platforms such as Douyin, TikTok,… | 65 | 1059 | active |
| meowcateatrat/elephant Elephant is an add-on for Free Download Manager that adds support for downloading videos from various websites, powered by YT-DLP. It is di… | 98 | 1058 | active |
| chainreactors/spray Spray is a high-performance HTTP directory fuzzing and content discovery tool written in Go, positioned as a next-generation alternative to… | 93 | 1058 | active |
| Cinvin/myuserscripts A Tampermonkey userscript for NetEase Cloud Music's web player that adds song downloading, cloud-disk transfer, fast cloud-disk upload, and… | 76 | 1058 | active |
| Rain120/qq-music-api A QQ Music API service built with Koa2 and TypeScript that proxies web-side QQ Music endpoints for songs, artists, playlists, and rankings.… | 76 | 1056 | active |
| lennybase/browsernode Browsernode is a TypeScript implementation of Browser-use that lets LLM-powered AI agents control a web browser via Playwright. It provides… | 34 | 1055 | active |
| Casvt/Kapowarr Kapowarr is a self-hosted web application for building and managing a digital comic book library, designed to fit into the *arr suite of me… | 87 | 1054 | active |
| kort0881/telegram-proxy-collector A Python CLI tool that automatically collects, analyzes, and filters MTProto and SOCKS5 proxies for Telegram. It decodes proxy secrets to d… | 79 | 1054 | active |
| Zarcolio/sitedorks A Python CLI tool that runs Google dork-style searches across multiple search engines (Google, Bing, DuckDuckGo, Yandex, Yahoo, Ecosia, Bra… | 76 | 1053 | active |
| Cloxl/xhshow A pure-algorithm Python library that generates Xiaohongshu (XHS/RedNote) request signature headers such as x-s, x-s-common, x-t, and x-rap-… | 76 | 1052 | active |
| dsclca12/auto_reg Any Auto Register is a self-hosted multi-platform account automatic registration and management system built with Python and Node.js. It in… | 52 | 1052 | active |
| tombcato/clash-ip-checker A Python automation tool for Clash proxy users that iterates through proxy nodes, checks each node's IP purity, bot ratio, and IP type via … | 44 | 1052 | active |
| robotshell/magicRecon MagicRecon is a Bash shell script that automates reconnaissance and vulnerability scanning of target domains, including subdomain enumerati… | 23 | 1052 | active |
| meetDeveloper/freeDictionaryAPI A free REST API that returns dictionary data for English words, including definitions, phonetics, audio pronunciations, origins, synonyms, … | 32 | 3590 | maintenance |
| Whisparr/Whisparr Whisparr is an adult movie collection manager for Usenet and BitTorrent users, forked from the Sonarr/Radarr family. It monitors RSS feeds … | 100 | 1050 | active |
| Henryhaohao/Bilibili_video_download A Python tool for downloading videos from Bilibili, supporting single-part and multi-part (分P) videos, bangumi episodes, and multiple downl… | 32 | 3583 | maintenance |
| jaimeiniesta/metainspector MetaInspector is a Ruby gem for web scraping that fetches a given URL and exposes its title, meta description, keywords, links, images, cha… | 69 | 1049 | active |
| lijiejie/GitHack GitHack is a Python CLI exploit tool that reconstructs a website's source code from an exposed .git folder. It parses the .git/index file, … | 32 | 3576 | maintenance |
| hanFengSan/eHunter eHunter is a Tampermonkey/userscript that injects a Vue 3-based comic reader UI into supported comic sites (EH/EXHentai, NHentai), offering… | 76 | 1045 | active |
| caolvchong-top/twitter_download A Python command-line tool that scrapes and downloads images, videos (including GIFs), and text from Twitter/X user timelines. It supports … | 75 | 1045 | active |
| mandatoryprogrammer/thermoptic Thermoptic is a stealth HTTP proxy that routes requests from any HTTP client (like curl) through a containerized Chrome instance so the tra… | 51 | 1045 | active |
| stay-leave/weibo-public-opinion-analysis A Python project for Weibo public opinion analysis that combines a web crawler, LDA topic modeling, sentiment analysis, and spatiotemporal … | 32 | 1045 | active |
| turicas/brasil.io The backend of Brasil.IO, a platform that collects, cleans, and publishes Brazilian public open datasets in accessible formats. It automate… | 77 | 1044 | active |
| tmwgsicp/wechat-download-api An open-source API service for fetching WeChat official account articles, generating standard RSS 2.0 feeds, and exporting entire account a… | 77 | 1043 | active |
| kelvinBen/AppInfoScanner A Python-based static information-gathering scanner for mobile apps (Android APK/DEX, iOS IPA/Mach-O) and static web content (HTML, JS, H5)… | 23 | 3554 | maintenance |
| HG-ha/ICP_Query A self-hosted service (Python and Rust implementations) that queries China MIIT ICP filing records for domains, apps, mini-programs, quick … | 95 | 1042 | active |
| TheAlgorithms/website The official website for The Algorithms, a static Next.js site that scrapes algorithm implementations from TheAlgorithms GitHub repositorie… | 77 | 1042 | active |
| buffer/thug Thug is a Python low-interaction honeyclient that mimics the behavior of a web browser to detect and emulate malicious web content. It comp… | 87 | 1041 | active |
| lm-rebooter/NuggetsBooklet A Node.js script that downloads Juejin (掘金) booklet content for personal study by reusing the user's authenticated browser cookies. It expo… | 66 | 1041 | active |
| badlogic/heissepreise A self-hostable grocery price search app that daily scrapes product and price data from major Austrian supermarket chains (Billa, Spar, Hof… | 60 | 1040 | active |
| igrigorik/ga-beacon A small Go service that acts as a collector-as-a-service for Google Analytics via the Measurement Protocol, letting you track page views wi… | 32 | 3541 | maintenance |
| GiantappMan/livewallpaper Giantapp Livewallpaper is an open-source wallpaper application for Windows 10/11 that supports both dynamic (video/animated) and static wal… | 74 | 1039 | active |
| Anorov/cloudflare-scrape A Python module (cfscrape) built on Requests that bypasses Cloudflare's JavaScript anti-bot challenge page ('I'm Under Attack Mode') so scr… | 23 | 3538 | maintenance |
| webwhiz-ai/webwhiz WebWhiz is an open-source, self-hostable application that trains a ChatGPT-powered chatbot on your website data by crawling your pages and … | 48 | 1037 | active |
| TheRook/subbrute SubBrute is a Python DNS meta-query spider that enumerates subdomains and arbitrary DNS record types by leveraging open resolvers to bypass… | 23 | 3526 | maintenance |
| Diving-Fish/maimaidx-prober A score tracker (prober) for the arcade rhythm game maimai DX that imports play records via a proxy tool and displays DX Rating and song sc… | 91 | 1032 | active |
| EdgeSecurityTeam/EHole EHole (棱洞) is a Go-based fingerprint identification tool for red team reconnaissance that pinpoints high-value, easily attackable systems (… | 23 | 3511 | maintenance |
| maxzhang666/OneKeyVip A multi-function browser userscript (compatible with Tampermonkey and ScriptCat) that bundles VIP video/music parsing, Bilibili cover fetch… | 76 | 1030 | active |
| Wikidepia/InstaFix InstaFix is a Go web service that serves fixed Instagram image and video embeds for Discord and Telegram by rewriting URLs (e.g., adding 'd… | 10 | 1029 | active |
| carcabot/tiktok-signature A self-hosted Node.js service that generates valid X-Bogus and X-Gnarly signature tokens for TikTok API requests using a headless browser r… | 93 | 1025 | active |
| wnma3mz/wechat_articles_spider A Python library for scraping WeChat Official Account articles, including article URLs, reading counts, likes, and comments. It can also do… | 23 | 3481 | maintenance |
| fossology/fossology FOSSology is an open source license compliance system and toolkit that scans software for licenses, copyrights, and export control data. It… | 87 | 1024 | active |
| agregarr/agregarr Agregarr is a self-hosted, Docker-based Plex Collections manager that automatically creates and refreshes collections from sources like Tra… | 67 | 1023 | active |
| am-will/codex-skills A collection of Codex/agent skills written in Shell covering planning, documentation access, prompting, frontend design guidance, Codex too… | 57 | 1023 | active |
| daijro/hrequests hrequests is a Python HTTP client library that replaces the requests library with browser TLS fingerprint replication, HTTP/2 support, and … | 30 | 1023 | active |
| hrithikkoduri/WebRover WebRover is an autonomous AI web agent that interprets user input, navigates websites via browser automation, and performs tasks or deep re… | 14 | 1022 | active |
| JosephLai241/URS URS (Universal Reddit Scraper) is a comprehensive command-line tool written in Python (with Rust components) for scraping and archiving Red… | 64 | 1020 | active |
| owner888/phpspider phpspider is a PHP web crawling framework that lets developers build scrapers with a simple config array, handling multi-process workers, l… | 23 | 3462 | maintenance |
| mxrch/GitFive GitFive is a Python-based OSINT CLI tool for investigating GitHub user profiles. It uncovers usernames, name history, email addresses, and … | 43 | 1019 | active |
| wujunwei928/parse-video A Go library and CLI tool that parses short-video share links from 25+ Chinese platforms (Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo, e… | 81 | 1018 | active |
| Vinyzu/Botright Botright is a Python browser automation framework built on Playwright that provides undetectable, fingerprint-changing stealth browsing. It… | 76 | 1017 | active |
| vasani-arpit/WBOT WBOT is a Node.js-based bot for WhatsApp Web that automates message replies using Puppeteer to control a browser. It is configurable via a … | 56 | 1013 | active |
| arabcoders/ytptube YTPTube is a self-hosted web-based download manager and automation layer for yt-dlp. It combines scheduled tasks, metadata-driven condition… | 90 | 1012 | active |
| davis7dotsh/my-pi-setup An opinionated configuration and extension setup for the Pi coding agent, adding themes, background terminals, subagents, workflows, and se… | 57 | 1012 | active |
| instagram4j/instagram4j instagram4j is an object-oriented, reverse-engineered Instagram Private API client library for Java. It lets developers log in, post, messa… | 81 | 1010 | active |
| pea3nut/Pxer Pxer is a userscript (installed via Tampermonkey) that acts as a crawler for pixiv.net, letting users batch-fetch artworks, collections, an… | 27 | 1010 | active |
| vimcolorschemes/vimcolorschemes A website and open-source app for browsing and discovering Vim and Neovim colorschemes from GitHub repositories. It tracks thousands of rep… | 97 | 1007 | active |
| wreq wreq is an ergonomic, privacy-aware HTTP client written in Rust with a Python binding (wreq-python) that provides high-fidelity browser TLS… | 90 | 1006 | active |
| ranahaani/GNews GNews is a lightweight Python package that queries the Google News RSS feed and returns article results as usable JSON. It supports keyword… | 68 | 1006 | active |
| Kappaemme-git/codex-first-customer-finder-skill A Codex skill (plugin) that takes a startup URL or product idea and produces an evidence-backed shortlist of potential first customers from… | 57 | 1006 | active |
| JoMingyu/google-play-scraper A Python library that provides APIs to crawl the Google Play Store for app details, reviews, and other data without any external dependenci… | 32 | 1006 | active |
| ma6254/FictionDown FictionDown is a Go-based command-line tool for batch downloading and crawling web novels from sites like Qidian and Biquge. It supports mu… | 24 | 1006 | active |
| shy1132/VacuumTube VacuumTube is an unofficial Electron-based desktop wrapper of YouTube Leanback, the official YouTube interface for consoles and Smart TVs. … | 86 | 1005 | active |
| ki9mu/ARL-plus-docker A Docker-based fork of ARL (Asset Reconnaissance Lighthouse) v2.6.2 that performs automated asset discovery and vulnerability scanning for … | 35 | 1005 | active |
| TypNull/Tubifarry Tubifarry is a Lidarr plugin that adds extra music sources to your library management, using Spotify as an indexer and YouTube or Soulseek … | 87 | 1003 | active |
| jackwener/wechat-article-to-markdown A Python CLI tool that fetches WeChat Official Account articles using anti-detection browser automation (Camoufox) and converts them to cle… | 48 | 1002 | active |
| gwen001/pentest-tools A collection of small custom security scripts in Bash, Python, and PHP for penetration testing and bug bounty quick tasks, covering DNS enu… | 23 | 3323 | maintenance |
| ping/instagram_private_api A Python client library for Instagram's private (undocumented) app and web APIs, with no third-party dependencies. It exposes app-only feat… | 23 | 3299 | maintenance |
| s-rah/onionscan OnionScan is a free and open source Go CLI tool for investigating Tor hidden services (.onion sites) on the Dark Web. It scans sites for op… | 23 | 3290 | maintenance |
| alixaxel/chrome-aws-lambda A TypeScript library that ships a size-optimized Chromium binary designed to run headless browser automation with Puppeteer (or Playwright)… | 32 | 3285 | maintenance |
| kevinzg/facebook-scraper A Python library for scraping public Facebook pages, groups, profiles, and posts without requiring an API key. It provides a simple get_pos… | 32 | 3270 | maintenance |
| archivy/archivy Archivy is a self-hostable personal knowledge repository that combines note-taking, bookmarking with full-page preservation, and a searchab… | 23 | 3268 | maintenance |
| cztomczak/cefpython CEF Python provides Python bindings for the Chromium Embedded Framework, letting Python applications embed a full Chromium-based browser. I… | 59 | 3237 | maintenance |
| wenbochang888/house A Java web scraper built with SpringBoot, HttpClient, and JSoup that crawls famous Tianya forum threads about China's housing market and co… | 32 | 3231 | maintenance |
| scrapy-plugins/scrapy-splash A Scrapy plugin that integrates the Splash headless browser service to enable crawling and scraping of JavaScript-rendered web pages. It pr… | 26 | 3227 | maintenance |
| ferventdesert/Hawk Hawk is a visual crawler and ETL IDE written in C#/WPF that lets users graphically scrape webpages, clean, transform, and store data withou… | 23 | 3213 | maintenance |
| atlas-comstock/NeteaseCloudMusicFlac A Python command-line script that downloads lossless FLAC music files to local storage based on a Netease Cloud Music playlist URL. It auto… | 32 | 3142 | maintenance |
| tomnomnom/httprobe httprobe is a Go command-line tool that takes a list of domains on stdin and probes for working HTTP and HTTPS servers, reporting which res… | 23 | 3119 | maintenance |
| CrawlScript/WebCollector WebCollector is an open-source Java web crawler framework that provides simple interfaces for building multi-threaded web crawlers quickly.… | 62 | 3083 | maintenance |
| 0x0be/yesitsme A Python CLI script for OSINT investigations that finds Instagram profiles matching a given name, e-mail, or phone number. It scrapes dumpo… | 32 | 3046 | maintenance |
| CodeRayZhang/Movie_Recommend A full-stack movie recommendation system built on Spark, including a Scrapy crawler, an SSM-based movie website, an admin backend, and a Sp… | 32 | 3007 | maintenance |
| mathieudutour/medium-to-own-blog A CLI tool that migrates a Medium blog to a self-hosted Gatsby-based blog in minutes. It exports Medium content into Markdown/MDX and scaff… | 32 | 3003 | maintenance |
| jaeles-project/gospider GoSpider is a fast web spider/crawler written in Go that crawls sites in parallel and extracts URLs from sitemaps, robots.txt, JavaScript f… | 23 | 2993 | maintenance |
| kotartemiy/newscatcher A Python package that programmatically collects normalized news articles from thousands of news websites, filterable by topic, country, and… | 32 | 2987 | maintenance |
| x0rz/tweets_analyzer A Python command-line script that scrapes a Twitter profile's tweets and analyzes metadata such as activity patterns, timezone, sources, ge… | 23 | 2981 | maintenance |