function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| hengliyin/cdfang-spider A full-stack web application that scrapes Chengdu Housing Association lottery housing data and presents it through interactive charts and s… | 53 | 1221 | active |
| warifp/FacebookToolkit A PHP command-line toolkit for retrieving Facebook account data via the Graph API, including access tokens, friend IDs, emails, and names, … | 23 | 1221 | active |
| cnbattle/douyin A Go-based crawler that scrapes Douyin (TikTok China) recommendation and search page video lists by controlling the mobile app on a real de… | 61 | 1219 | active |
| Silent1566/OmniBox-Spider A collection of spider (scraper) sources and interfaces for the OmniBox media application, aggregated from publicly available internet info… | 58 | 1213 | active |
| Sak32009/GetDataFromSteam-SteamDB A userscript that extracts game and DLC data from Steam store pages and SteamDB. It runs via userscript managers like Tampermonkey or Viole… | 76 | 1212 | active |
| taranis-ai/taranis-ai Taranis AI is a self-hosted open-source OSINT platform that collects news articles from web sources and uses NLP/AI to enrich, cluster, and… | 95 | 1207 | active |
| lexmount/moli Moli is a lightweight, fast headless browser built in Rust (on Servo technology) designed for AI agents to fetch, render, and extract web p… | 79 | 1206 | active |
| browserable/browserable Browserable is an open-source, self-hostable JavaScript library for building AI browser agents that navigate sites, fill out forms, click b… | 37 | 1206 | active |
| CERT-Polska/Artemis Artemis is a modular, open-source vulnerability scanner developed by CERT Polska that checks website security at scale. It automatically ge… | 74 | 1204 | active |
| gityuanbao/share A personal open-source repository whose main component is akshare_collector, a Python tool built on AKShare that collects Chinese financial… | 64 | 1201 | active |
| scrapinghub/splash Splash is a lightweight, scriptable headless browser exposed as a service with an HTTP API, implemented in Python 3 using Twisted and Qt5. … | 32 | 4187 | maintenance |
| ycdxsb/PocOrExp_in_Github A Python CLI tool that automatically aggregates proof-of-concept (POC) and exploit (EXP) code from GitHub by CVE ID, using CVE information … | 77 | 1198 | active |
| RyensX/MediaBox MediaBox is an Android 'universal media container' app that aggregates video, manga, and other media through a WeChat-mini-program-like plu… | 23 | 1198 | active |
| internetwache/GitTools GitTools is a collection of three shell/Python scripts for finding and exploiting websites that publicly expose their .git directory. It in… | 64 | 4178 | maintenance |
| niudaii/zpscan zpscan is a Go-based command-line information gathering and reconnaissance tool for security assessments. It bundles subdomain enumeration,… | 23 | 1196 | active |
| instantbox/instantbox instantbox is a self-hosted service that spins up temporary, clean Linux containers (Ubuntu, CentOS, Arch, Debian, Fedora, Alpine) with ins… | 32 | 4176 | maintenance |
| teaSummer/MCiSEE MCiSEE is a web application that aggregates and presents Minecraft resources (software, tools, websites) collected from the internet in an … | 69 | 1194 | active |
| CarrotRub/Fit-Launcher Fit Launcher is a desktop game launcher built with Rust, Tauri, and SolidJS for downloading and playing games from FitGirl Repacks. It uses… | 77 | 1193 | active |
| constverum/ProxyBroker ProxyBroker is an asynchronous Python tool that finds public HTTP(S) and SOCKS4/5 proxies from ~50 sources and concurrently checks their ty… | 32 | 4159 | maintenance |
| jayus0821/swagger-hack A Python CLI tool that automatically crawls all endpoints exposed by leaked Swagger/OpenAPI documentation and sends configured test request… | 57 | 1189 | active |
| xenova/chat-downloader Chat Downloader is a Python tool and library for retrieving chat messages from livestreams, videos, clips, and past broadcasts on platforms… | 45 | 1189 | active |
| intoli/user-agents A JavaScript/TypeScript npm package for generating random user agents weighted by real-world market share, with daily-updated data. It also… | 77 | 1188 | active |
| kmille/deezer-downloader A self-hosted Python service with a simple web frontend for downloading songs, albums, and playlists from Deezer (and via yt-dlp), with ID3… | 74 | 1186 | active |
| neon-mmd/websurfx Websurfx is an open-source meta search engine written in Rust that aggregates results from multiple search engines while respecting user pr… | 86 | 1184 | active |
| LainsNL/OutlookRegister A Python automation tool that mass-registers Outlook/Hotmail email accounts using browser automation (Playwright or Patchright) with simula… | 62 | 1184 | active |
| gooseworks-ai/goose-skills A library of 200+ AI agent skills and data APIs for growth and go-to-market work, installable into coding agents like Claude Code, Cursor, … | 58 | 1184 | active |
| goodreasonai/ScrapeServ ScrapeServ is a self-hosted API service that accepts a URL and returns the website's data along with browser screenshots, using Playwright … | 25 | 1181 | active |
| shobrook/rebound Rebound is a command-line tool that runs your file and, when an exception is thrown, instantly fetches related Stack Overflow questions and… | 23 | 4114 | maintenance |
| austin-weeks/miasma Miasma is a lightweight Rust web server that traps AI web scrapers in an endless pit of poisoned training data and self-referential links. … | 82 | 1179 | active |
| ycngmn/Nobook Nobook is a lightweight, ad-free Android client for browsing Facebook, built with Kotlin and Jetpack Compose on top of a WebView. It blocks… | 10 | 1179 | active |
| rchipka/node-osmosis Osmosis is an HTML/XML parser and web scraper library for Node.js built on native libxml C bindings. It offers a chainable, promise-like in… | 32 | 4107 | maintenance |
| grangier/python-goose Python-Goose is a Python library that extracts the main body text, metadata, top image, and embedded videos from news article web pages. It… | 64 | 4106 | maintenance |
| scdl-org/scdl scdl is a Python command-line tool for downloading music, playlists, likes, and reposts from SoundCloud, with automatic ID3 metadata taggin… | 67 | 4099 | maintenance |
| N0rz3/Phunter Phunter is a Python CLI OSINT tool that gathers information about phone numbers, including operator, line type, location, reputation, spam … | 26 | 1175 | active |
| apache/groovy-geb Apache Geb is a Groovy-based browser automation library built on WebDriver, combining jQuery-like content selection with Page Object modell… | 76 | 1173 | active |
| myfanhua/turb-gpt-free-register A Python tool that bulk-registers ChatGPT/OpenAI accounts using pure-protocol requests or anti-fingerprint browser automation (RoxyBrowser,… | 58 | 1172 | active |
| Evolution0/bandcamp-dl A Python command-line tool for downloading albums and tracks from bandcamp.com, with options for filename templating, album art embedding, … | 65 | 1171 | active |
| online-judge-tools/oj A command-line tool that automates solving problems on online judges like AtCoder, Codeforces, and HackerRank. It downloads sample and syst… | 23 | 1169 | active |
| rverton/webanalyze webanalyze is a Go port of Wappalyzer that detects the technologies used on websites, built for performant mass scanning of large host list… | 71 | 1168 | active |
| aooiuu/any-reader Any-Reader is an open-source, cross-platform content aggregation tool that unifies novels, manga, video, and audio from user-defined source… | 73 | 1167 | active |
| xeco23/WasIstLos WasIstLos is an unofficial WhatsApp desktop client for Linux, written in C++ using gtkmm and WebKitGTK to wrap WhatsApp Web. It adds deskto… | 10 | 1166 | active |
| asdfghj1237890/WebVideo2NAS A self-hosted pipeline consisting of a Chrome extension and a Dockerized FastAPI backend that detects HLS (M3U8), DASH (MPD), MP4, and MOV … | 80 | 1161 | active |
| CMHopeSunshine/LittlePaimon LittlePaimon is a multifunctional Genshin Impact chatbot built on the NoneBot2 framework, supporting the OneBot protocol for QQ. It provide… | 27 | 1160 | active |
| zu1k/proxypool A Go service that automatically crawls proxy nodes (ss, ssr, vmess, trojan) from Telegram channels, subscription URLs, and the public inter… | 23 | 4027 | maintenance |
| 0x727/ShuiZe_0x727 ShuiZe_0x727 is a Python-based automated information gathering (reconnaissance) tool for red team operators. Given a root domain, C-segment… | 23 | 4019 | maintenance |
| fanpei91/torsniff torsniff is a Go CLI tool that sniffs torrent metadata from the BitTorrent network by participating in the DHT and connecting to peers to d… | 10 | 4014 | maintenance |
| librariesio/libraries.io Libraries.io is an open source discovery service that indexes millions of packages across 32 package managers, letting developers search an… | 75 | 1156 | active |
| evyatarmeged/Raccoon Raccoon is a Python-based offensive security CLI tool for reconnaissance and information gathering. It performs DNS lookups, WHOIS, TLS ana… | 67 | 4001 | maintenance |
| EmilStenstrom/justhtml JustHTML is a pure Python HTML5 parser with browser-style error recovery, safe-by-default sanitization, CSS selector querying, and serializ… | 83 | 1151 | active |
| agentbay-ai/wuying-agentbay-sdk Multi-language SDK (Python, TypeScript, Go, Java) for Wuying AgentBay, Alibaba Cloud's cloud sandbox platform built for AI agents. It lets … | 76 | 1148 | active |
| filipedeschamps/rss-feed-emitter A Node.js library that aggregates RSS and Atom news feeds and emits events for every new item published. It automatically manages feed hist… | 72 | 1147 | stable |
| yunginnanet/HellPot HellPot is a cross-platform HTTP honeypot that punishes bots ignoring robots.txt by streaming an infinite Markov-chain-generated page of ps… | 49 | 1147 | active |
| random-robbie/My-Shodan-Scripts A collection of Python 3 scripts for querying the Shodan search engine to find exposed devices and services on the internet. It bundles man… | 60 | 1146 | active |
| 0xsha/CloudBrute CloudBrute is a Go CLI tool that enumerates a company's infrastructure, files, and applications across major cloud providers (Amazon, Googl… | 27 | 1145 | active |
| datawhores/OF-Scraper OF-Scraper is a command-line tool for downloading media from OnlyFans and performing bulk actions like liking or unliking posts. It is a re… | 80 | 1143 | active |
| nicolomantini/LinkedIn-Easy-Apply-Bot A Python bot that automates applying to jobs on LinkedIn using the Easy Apply feature. It reads search preferences, resume uploads, and bla… | 57 | 1143 | active |
| Tucsky/aggr AGGR (SignificantTrade) is a Vue.js web application that aggregates live cryptocurrency trades from many exchanges (Binance, Coinbase, BitM… | 59 | 1142 | active |
| AndyTheFactory/newspaper4k Newspaper4k is a Python library and CLI for scraping and curating news articles, extracting text, titles, authors, publish dates, and metad… | 84 | 1140 | active |
| kevthehermit/PasteHunter PasteHunter is a Python 3 application that queries public pastebin-style sites (pastebin.com, GitHub gists, slexy, stackexchange, etc.) and… | 50 | 1137 | active |
| ChineseSubFinder/ChineseSubFinder A self-hosted Go application that automatically downloads Chinese subtitles for movies and TV shows from subtitle websites. It integrates w… | 23 | 3932 | maintenance |
| fwonggh/Bthub Bthub is a magnet link and torrent search engine, and this repository serves as its official address release page listing current and backu… | 76 | 1133 | active |
| hasanfirnas/symbiote Symbiote is a Python-based social engineering tool that generates a phishing page to trick a target into granting camera permission, then c… | 37 | 1132 | active |
| auto-novel/auto-novel AutoNovel is a website application that automatically machine-translates light novels (web novels, bunko novels, and local files) using LLM… | 77 | 1130 | active |
| rebane2001/xikipedia Xikipedia is a web application that presents Wikipedia content as a social media feed, using a simple non-ML algorithm that learns user eng… | 53 | 1129 | active |
| codelibs/fess Fess is an open-source, self-hosted enterprise search server built on OpenSearch with a browser-based admin UI, built-in crawlers for web s… | 95 | 1128 | active |
| johnwmillr/LyricsGenius A Python client library that wraps the Genius.com API and scrapes song lyrics, artist, and album metadata. It provides simple search and do… | 69 | 1121 | active |
| QIN2DIM/epic-awesome-gamer A Python application that automatically claims weekly free games and monthly content from the Epic Games Store. It uses browser automation … | 56 | 1121 | active |
| webrecorder/browsertrix-crawler Browsertrix Crawler is a standalone browser-based high-fidelity web crawling system that runs in a single Docker container. It uses Puppete… | 99 | 1120 | active |
| Haleydu/Cimoc Cimoc is an open-source Android manga/comic reader app that aggregates multiple online comic sources. It supports page and scroll reading m… | 23 | 3855 | maintenance |
| r3nt0n/bopscrk bopscrk is a Python CLI tool that generates smart, targeted wordlists for password cracking, combining user-provided words with transformat… | 23 | 1117 | active |
| Doriandarko/make-it-heavy A Python framework that emulates Grok Heavy-style deep analysis by orchestrating multiple specialized AI agents in parallel via OpenRouter'… | 32 | 1116 | active |
| shanmiteko/LotteryAutoScript A Node.js automation script that monitors Bilibili dynamic posts for lottery/giveaway draws and automatically participates by liking, comme… | 89 | 1115 | active |
| platonai/Browser4 Browser4 is an AI-native browser engine built in Kotlin for autonomous agents, intelligent data extraction, and large-scale web automation.… | 100 | 1114 | active |
| elixir-crawly/crawly Crawly is a high-level web crawling and scraping framework for Elixir, modeled after Scrapy, where developers define spiders that fetch pag… | 37 | 1114 | active |
| nottelabs/reverse-api-engineer Reverse API Engineer is a Python CLI tool that captures browser network traffic (HAR) from a website and uses a configured AI model to gene… | 84 | 1113 | active |
| gautamkrishnar/socli SoCLI is a Python command-line client for Stack Overflow that lets developers search and browse questions and answers without leaving the t… | 56 | 1111 | active |
| bellingcat/auto-archiver A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supp… | 92 | 1109 | active |
| mrkrsl/web-search-mcp A locally hosted MCP (Model Context Protocol) server written in TypeScript that gives local LLMs web search capabilities without requiring … | 36 | 1109 | active |
| gxtrobot/bustag Bustag is a self-hosted web application that periodically scrapes new media title (番号) listings, lets users tag items as liked or disliked,… | 23 | 3814 | maintenance |
| leovan/SciHubEVA SciHubEVA is a cross-platform GUI application for searching and downloading papers from Sci-Hub, built with Python and Qt (PySide6/QML). It… | 74 | 1107 | active |
| Komet/MediaElch MediaElch is a cross-platform desktop media manager for Kodi that scrapes and organizes metadata for movies, TV shows, concerts, and music.… | 67 | 1102 | active |
| vifreefly/kimuraframework Kimuraframework (Kimurai) is a Ruby web scraping framework with an AI-assisted DSL: an LLM generates XPath selectors from a schema on first… | 61 | 1102 | active |
| GPTaku Plugins insane-search is a Claude Code plugin that reads public web pages that would otherwise be blocked (403, CAPTCHA, WAF), escalating through p… | 60 | 1102 | active |
| alvarorichard/GoAnime GoAnime is a terminal-based (TUI) anime browser written in Go that lets users search, stream, and download anime episodes directly in mpv. … | 94 | 1101 | active |
| requireCool/stealth.min.js A repository that automatically generates and publishes the newest stealth.min.js file every Monday at 6 AM UTC. The file is a drop-in Java… | 77 | 1099 | active |
| iFurySt/RedNote-MCP A Model Context Protocol (MCP) server that lets AI clients access RedNote (Xiaohongshu) content, including keyword search and note retrieva… | 19 | 1097 | active |
| SteamTracking/SteamTracking A project that tracks and reverse-engineers Steam and Valve data, including protobuf definitions and changes across Steam services. It moni… | 77 | 1095 | active |
| projectdiscovery/wappalyzergo A high-performance Go library that ports the Wappalyzer technology detection stack, identifying web technologies from HTTP headers and HTML… | 95 | 1092 | active |
| eshaham/israeli-bank-scrapers A TypeScript library providing scrapers for all major Israeli banks and credit card companies, published as the npm package israeli-bank-sc… | 96 | 1091 | active |
| Tuhinshubhra/RED_HAWK RED_HAWK is a PHP-based all-in-one reconnaissance and vulnerability scanning tool for websites. It performs information gathering (whois, D… | 32 | 3748 | maintenance |
| YaoZeyuan/stablog Stablog (稳部落) is a desktop application that backs up and exports a user's Weibo posts into searchable HTML and PDF ebooks. It logs into the… | 66 | 1087 | active |
| GeneralMills/pytrends Pytrends is an unofficial Python library providing a pseudo API for Google Trends, enabling automated downloading of trend reports. It wrap… | 10 | 3730 | maintenance |
| JimmXinu/FanFicFare FanFicFare is a Python tool that downloads stories from over 100 fanfiction and web fiction sites and converts them into EPUB (and HTML) eB… | 98 | 1085 | active |
| sharebook-kr/pykrx PyKrx is a Python library that scrapes stock and bond market data from the Korea Exchange (KRX) and Naver. It provides APIs for querying ti… | 76 | 1084 | active |
| Tyrrrz/YoutubeExplode A .NET library providing an abstraction layer over YouTube's internal API to query metadata for videos, playlists, and channels, and to res… | 99 | 3717 | maintenance |
| ruipgil/scraperjs Scraperjs is a Node.js web scraping library offering two scrapers: a lightweight StaticScraper using cheerio for static HTML, and a Dynamic… | 32 | 3714 | maintenance |
| techtanic/Discounted-Udemy-Course-Enroller A Python application (with GUI and CLI variants) that scrapes websites for 100% off Udemy course coupons and automatically enrolls the user… | 71 | 1081 | active |
| nasa/apod-api NASA's open-source microservice that serves the Astronomy Picture of the Day (APOD) API, returning JSON metadata and image links parsed fro… | 65 | 1080 | active |
| viu-media/viu Viu is a terminal-based anime client that provides a rich TUI for browsing, searching, and managing your AniList library, along with stream… | 10 | 1079 | active |
| m-sec-org/EZ EZ is a cross-platform vulnerability scanner that combines information gathering, port scanning, service brute-forcing, URL crawling, finge… | 24 | 1078 | active |