function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| ssili126/tv A Python tool that automatically collects IPv4 hotel IPTV live stream sources (including CCTV, satellite, and some local Chinese channels),… | 70 | 1956 | active |
| vmoranv/jshookmcp An MCP server exposing 600+ tools across 34 domains for JavaScript reverse engineering and security research, including browser automation,… | 59 | 1954 | active |
| AWeirdDev/flights fast-flights is a Python library that scrapes Google Flights by generating Base64-encoded Protobuf query strings, returning strongly-typed … | 86 | 1939 | active |
| feder-cr/invisible_playwright A Python library that provides an antidetect, stealth-patched Firefox build for Playwright, with fingerprints set at the C++ engine level a… | 81 | 1938 | active |
| linbailo/zyqinglong A collection of self-use scripts for the Qinglong panel that automates daily check-ins and coupon collection for Chinese apps like Didi, Me… | 40 | 1925 | active |
| damoeb/rss-proxy RSS-proxy is a self-hostable web service that generates RSS, ATOM, or JSON feeds from almost any static website by analyzing its HTML struc… | 32 | 1924 | active |
| SleepingBag945/dddd dddd is a Go-based batch information gathering and supply-chain vulnerability detection CLI tool designed to streamline red team workflows.… | 19 | 1924 | active |
| MarginaliaSearch/MarginaliaSearch Marginalia Search is an independent, open-source internet search engine that indexes text-oriented, non-commercial, small and old websites.… | 65 | 1923 | active |
| ZeroPointSix/outlookEmailPlus OutlookMail Plus is a self-hosted email manager purpose-built for account registration and verification workflows. It automates fetching ve… | 81 | 1917 | active |
| KEV0143/Parser-Chitai-Gorod A Python-based scraper for the Russian online bookstore Chitai-Gorod that collects book URLs across catalog pages and extracts structured p… | 29 | 1914 | active |
| aeonfun/opendia OpenDia is an open-source browser extension (npm package) that connects your Chrome, Arc, or Firefox browser to AI models via the Model Con… | 83 | 1913 | active |
| lanyeeee/bilibili-video-downloader A cross-platform GUI desktop application built with Tauri for downloading videos, audio, subtitles, danmaku, and covers from Bilibili. It s… | 75 | 1912 | active |
| extractus/article-extractor A TypeScript library that extracts the main article content, title, image, and metadata from a given URL or raw HTML string. It supports cu… | 98 | 1909 | active |
| sjdonado/idonthavespotify A web app that converts music links between streaming services like Spotify, Apple Music, YouTube Music, Tidal, and Deezer. It parses the s… | 75 | 1908 | active |
| karlicoss/promnesia Promnesia is a browser extension plus Python backend that enhances browsing history by annotating visited pages with context from multiple … | 82 | 1895 | active |
| bellingcat/octosuite Octosuite is a terminal-based toolkit for analyzing GitHub data, usable as an interactive TUI, a CLI, or a Python library. It queries user,… | 76 | 1895 | active |
| trevorhobenshield/twitter-api-client A Python library implementing X/Twitter's v1, v2, and GraphQL APIs for automation and scraping. It supports account actions like tweeting, … | 30 | 1894 | active |
| xtekky/TikTok-ViewBot A Python-based TikTok view bot that sends fake views via HTTP requests against Zefoy, without needing Selenium. It includes an automatic ca… | 35 | 1889 | active |
| nsonaniya2010/SubDomainizer SubDomainizer is a Python CLI tool that discovers hidden subdomains and secrets in webpages, external JavaScript files, GitHub, and local f… | 66 | 1886 | active |
| zema1/watchvuln WatchVuln is a self-hosted Go service that scrapes high-quality vulnerability sources (Aliyun AVD, Chaitin, OSCS, Qianxin TI, Seebug, CISA … | 62 | 1885 | active |
| scholarly-python-package/scholarly scholarly is a Python library for retrieving author and publication metadata from Google Scholar through a friendly, Pythonic API. It handl… | 56 | 1879 | active |
| degoog-org/degoog Degoog is a self-hosted search aggregator that queries multiple search engines in parallel from your own server and merges results into a s… | 80 | 1878 | active |
| 404-novel-project/novel-downloader An extensible userscript (Tampermonkey/Greasemonkey/Violentmonkey) that downloads novels from many Chinese web novel sites and exports them… | 76 | 1877 | active |
| coder-hxl/x-crawl x-crawl is a flexible Node.js crawler library that supports crawling dynamic pages, static pages, API data, and files, with optional AI ass… | 66 | 1877 | active |
| joyce677/TrendRadar TrendRadar is a lightweight self-hosted hot-topic monitor that aggregates trending news from 35 Chinese platforms (Weibo, Zhihu, Bilibili, … | 62 | 1877 | active |
| null2264/yokai Yōkai is a free and open source manga reader app for Android, forked from Tachiyomi/Mihon. It supports local and online manga reading with … | 93 | 1875 | active |
| ThePhaseless/Byparr Byparr is a self-hosted Python service that solves antibot browser challenges (like Cloudflare checks) and returns valid clearance cookies … | 90 | 1865 | active |
| 0xKayala/NucleiFuzzer NucleiFuzzer is a Python-based automation tool that combines URL discovery tools (ParamSpider, Waybackurls, Gauplus, Hakrawler, Katana) wit… | 66 | 1862 | active |
| tsingyuai/growth-lab Growth Lab is an open-source, end-to-end AI growth system that runs product marketing and user-acquisition loops through natural-language c… | 56 | 1861 | active |
| pmh1314520/WebRPA WebRPA is an open-source, no-code visual RPA tool for building automation workflows by dragging and connecting modules, covering web scrapi… | 82 | 1856 | active |
| ARC-MX/sgcc_electricity_new A Home Assistant integration (deployed via Docker) that scrapes China State Grid (SGCC) accounts to fetch electricity billing and usage dat… | 89 | 1855 | active |
| sarperavci/GoogleRecaptchaBypass A Python library that automatically solves Google reCAPTCHA v2 challenges in under five seconds using browser automation with DrissionPage … | 68 | 1855 | active |
| download-directory/download-directory.github.io A web app that lets users download a single subdirectory from a GitHub repository as a zip file, filling a gap in GitHub's native functiona… | 69 | 1852 | active |
| tiantianGPU/reg-factory A Python-based local web console that automates bulk registration of email and AI service accounts (Outlook, Gmail, ChatGPT, Grok, Claude, … | 80 | 1850 | active |
| ZianTT/BHYG BHYG is a script/tool for automatically grabbing tickets for Bilibili World (BW) events via Bilibili's member purchase (会员购) platform. It a… | 73 | 1849 | active |
| GuDaStudio/GrokSearch GrokSearch is an MCP server built on FastMCP that gives Claude Code and other LLM clients real-time web access via a dual-engine architectu… | 48 | 1849 | active |
| 1234567Yang/cf-proxy-ex A Cloudflare Workers-based super proxy that lets users access websites like GitHub, DuckDuckGo, and StackOverflow through a different URL w… | 69 | 1848 | active |
| wapiti-scanner/wapiti Wapiti is an open-source black-box web vulnerability scanner written in Python that crawls deployed web applications and fuzzes scripts and… | 98 | 1846 | active |
| mwmbl/mwmbl Mwmbl is an open source, non-profit web search engine with no ads or tracking, where the community determines rankings and runs distributed… | 77 | 1844 | active |
| AnySearch AnySearch is a unified real-time search engine service for AI agents, distributed as an agent skill package and an MCP server. It provides … | 58 | 1838 | active |
| abinthomasonline/repo2txt A browser-based tool that converts GitHub, GitLab, Azure DevOps repositories, local folders, or ZIP files into a single formatted text file… | 65 | 1835 | active |
| microlinkhq/browserless A Node.js library that wraps Puppeteer to provide a production-ready headless Chrome/Chromium driver with built-in screenshot, PDF generati… | 95 | 1831 | active |
| 78778443/QingScan QingScan is a self-hosted, open-source security operations platform that unifies vulnerability scanning, code auditing, asset inventory, an… | 66 | 1831 | active |
| tryolabs/requestium Requestium is a Python library that merges Requests, Selenium, and Parsel into a single integrated tool for web automation. It lets scripts… | 77 | 1830 | active |
| lds133/weather_landscape A Python application that renders weather forecasts as a stylized landscape image instead of numeric dashboards, encoding time, temperature… | 56 | 1827 | active |
| initstring/linkedin2username A Python OSINT tool that scrapes LinkedIn employee lists for a target company and generates multiple probable username formats (e.g., first… | 76 | 1825 | active |
| 1N3/BlackWidow BlackWidow is a Python-based web application spider that crawls a target site to collect URLs, dynamic parameters, subdomains, email addres… | 57 | 1821 | active |
| autoclaw-cc/xiaohongshu-skills A set of AI agent skills (SKILL.md format) plus a Chrome extension that automates Xiaohongshu (RED) using your real logged-in browser sessi… | 70 | 1819 | active |
| petronny/gfwlist2pac A tool that automatically converts the gfwlist proxy rules into a PAC (Proxy Auto-Config) file every day. The generated gfwlist.pac is serv… | 77 | 1816 | active |
| enetx/surf Surf is an advanced HTTP client library for Go with fluent, chainable API design. It supports browser impersonation (Chrome/Firefox), JA3/J… | 83 | 1808 | active |
| POf-L/Fanqie-novel-Downloader A cross-platform desktop and mobile application built with Rust and Tauri v2 for searching, reading, and downloading novels from Fanqie Nov… | 82 | 1801 | active |
| LagradOst/QuickNovel QuickNovel is a free, ad-free, open-source Android app for downloading novels from many web novel sites, which also functions as an EPUB re… | 99 | 1786 | active |
| emacs-elfeed/elfeed Elfeed is an extensible web feeds client for Emacs supporting Atom, RSS, and JSON Feed formats. It provides a search-based UI inspired by n… | 67 | 1784 | active |
| kost/dvcs-ripper dvcs-ripper is a set of Perl command-line tools that download (rip) web-accessible version control repositories such as GIT, SVN, Mercurial… | 32 | 1784 | stable |
| Masterminds/html5-php A standards-compliant HTML5 parser and serializer written entirely in PHP. It parses HTML5 documents and fragments into standard PHP DOM ob… | 87 | 1782 | stable |
| mdc-ng/mdc-ng A self-hosted media metadata scraper and organizer for adult video libraries, written in Rust with a Next.js web UI. It scrapes metadata fr… | 83 | 1782 | active |
| ShunCai/QZoneExport A browser extension (Chrome/Edge/Firefox, Manifest V3) that backs up QQ Zone data—posts, blogs, private diaries, albums, videos, comments, … | 88 | 1780 | active |
| ScrapeCreators/social-media-research-skills A collection of AI agent skills for social media research built on the ScrapeCreators scraping API. The skills give agents complete workflo… | 58 | 1777 | active |
| bulianglin/psub psub is a proxy subscription conversion tool deployed on Cloudflare Workers that acts as a reverse proxy for subscription conversion backen… | 10 | 1772 | active |
| inulute/medium-unlocker Medium Unlocker is an Android app (with an accompanying web frontend) that bypasses Medium's paywall by redirecting article URLs to the fre… | 91 | 1767 | active |
| deweizhu/bookget bookget is a Go-based command-line tool for downloading digitized ancient books and rare texts from 50+ digital libraries. It ships prebuil… | 60 | 1764 | active |
| cambecc/earth Earth is a JavaScript web application that visualizes global weather, wind, ocean currents, and related atmospheric conditions on an animat… | 32 | 6588 | maintenance |
| LoseNine/ruyipage RuyiPage is a Python browser automation framework built on Firefox and the WebDriver BiDi protocol, shipping with an anti-detection Firefox… | 79 | 1759 | active |
| josh0xA/darkdump Darkdump is an open-source OSINT tool for querying multiple dark web search engines and scraping onion site results for emails, metadata, k… | 70 | 1757 | active |
| website-scraper/node-website-scraper A Node.js library that downloads entire websites to a local directory, including HTML, CSS, images, and JavaScript assets. It parses HTTP r… | 72 | 1751 | active |
| adsbypasser/adsbypasser AdsBypasser is a lightweight userscript that automatically skips countdown ads, continue/redirect pages, and prevents ad pop-up windows acr… | 99 | 1750 | active |
| Aas-ee/open-webSearch Open-WebSearch is a TypeScript tool providing an MCP server, CLI, and local daemon for multi-engine web search and content retrieval withou… | 79 | 1747 | active |
| sindresorhus/pageres-cli A Node.js command-line tool that captures screenshots of websites at multiple resolutions using headless Chrome (Puppeteer). It is useful f… | 44 | 1743 | active |
| IvanGlinkin/Fast-Google-Dorks-Scan A shell-based OSINT tool that automates Google dork searches against a target website to uncover admin panels, exposed file types, and path… | 46 | 1742 | active |
| egoist/sitefetch A Node.js CLI tool that crawls an entire website and saves its pages as a single text file, using Mozilla Readability to extract clean cont… | 22 | 1736 | active |
| dilame/instagram-private-api A NodeJS/TypeScript SDK providing a client for Instagram's private (undocumented) API, enabling full programmatic access to feeds, direct m… | 23 | 6469 | maintenance |
| cxOrz/chaoxing-signin A Node.js/TypeScript tool that automates sign-in for the Chaoxing (Superstar Learning) online course platform, supporting normal, photo, ge… | 10 | 1724 | active |
| hzm0321/real-time-fund A Next.js web application for real-time mutual fund valuation and top-holdings stock tracking, primarily for Chinese funds, with portfolio,… | 78 | 1723 | active |
| claffin/cloudproxy CloudProxy is a self-hosted Python tool that provisions and manages proxy servers across multiple cloud providers, rotating IPs to improve … | 81 | 1722 | active |
| Jules-WinnfieldX/CyberDropDownloader A Python-based bulk downloader that scrapes and downloads files from Cyberdrop.me and dozens of other file hosts and image galleries. It is… | 10 | 1720 | active |
| jldbc/pybaseball pybaseball is a Python package that scrapes and retrieves current and historical baseball statistics from sources like MLB Statcast (Baseba… | 50 | 1719 | active |
| qinlili23333/ctfileGet A web-based resolver that obtains one-time direct download URLs for files hosted on Chengtong Network Disk (ctfile/城通网盘) using its official… | 55 | 1714 | active |
| hangone/WeBan WeBan is a Python-based automation tool that automatically completes courses and exams on the Weiban (安全微伴) university safety education pla… | 85 | 1712 | active |
| 3441293738/creatorhub CreatorHub is a self-hosted web panel built with Python and FastAPI for managing, monitoring, scraping, downloading, and publishing content… | 58 | 1712 | active |
| zstmfhy/zlibrary-to-notebooklm A Python CLI tool that automatically downloads books from Z-Library and uploads them to Google NotebookLM in one command. It uses Playwrigh… | 43 | 1711 | active |
| Python3Spiders/WeiboSuperSpider A Weibo (Chinese microblog) scraping toolbox in Python covering users, topics, and comments, with extras like image downloading, sentiment … | 75 | 1705 | active |
| aisingapore/TagUI TagUI is a free, open-source robotic process automation (RPA) tool from AI Singapore that lets users write simple text flows to automate re… | 65 | 6325 | maintenance |
| utkusen/urlhunter urlhunter is a Go-based recon CLI tool that searches URLs exposed via shortener services like bit.ly and goo.gl. It downloads daily URLTeam… | 24 | 1697 | active |
| oxylabs/google-play-scraper A free Python-based Google Play Store scraper that collects public app, movie, and book data via search queries. It is a companion tool to … | 67 | 1693 | active |
| openkursar/hello-halo Halo is an open-source desktop AI workstation that wraps frontier coding agents like Claude Code and Codex in a visual GUI, with a pluggabl… | 81 | 1688 | active |
| amaancoderx/npxskillui SkillUI is a Node.js CLI that crawls websites, git repos, or local codebases and extracts their complete design system (colors, typography,… | 51 | 1688 | active |
| LearnPrompt/ai-news-radar AI News Radar is an automated 24-hour AI/tech news aggregator that fetches sources, deduplicates stories, and scores headlines with three p… | 80 | 1687 | active |
| Ge0rg3/requests-ip-rotator A Python library that mounts AWS API Gateway as a proxy onto requests sessions, rotating source IPs on every request using AWS's large IP p… | 90 | 1675 | active |
| jarun/googler googler is a Python command-line tool for performing Google web, news, and video searches directly from the terminal. It displays titles, U… | 10 | 6204 | maintenance |
| vibheksoni/stealth-browser-mcp A Python MCP server that exposes stealth browser automation (via nodriver and Chrome DevTools Protocol) to AI agents, letting them navigate… | 61 | 1674 | active |
| rushter/selectolax Selectolax is a fast Python HTML5 parser library written in Cython, binding to the Modest and Lexbor parsing engines. It provides CSS selec… | 94 | 1665 | active |
| NeteaseCloudMusicApiEnhanced/api-enhanced A Node.js API service providing comprehensive access to NetEase Cloud Music (网易云音乐) endpoints, a half-refactored and enhanced fork of the p… | 84 | 1661 | active |
| ReScienceLab/opc-skills A collection of open-source Agent Skills (folders of instructions, scripts, and resources) designed for solopreneurs, indie hackers, and on… | 73 | 1658 | active |
| itteco/iframely Iframely is a self-hosted oEmbed proxy and URL metadata service that takes any URL and returns rich media embed codes and page metadata. It… | 67 | 1649 | active |
| nasa-gibs/worldview NASA Worldview is an interactive web application for browsing over 1000 global, full-resolution satellite imagery layers from NASA's Global… | 95 | 1648 | active |
| srx-2000/spider_collection A collection of Python web crawler scripts targeting sites like Bilibili, Zhihu, Weibo, NetEase Music, GitHub, and Anjuke, built with reque… | 32 | 1646 | active |
| Benexl/yt-x A POSIX-compliant shell script that lets you browse YouTube and other yt-dlp-supported sites from the terminal using fzf or from an app lau… | 85 | 1645 | active |
| itsToggle/plex_debrid A Python application that monitors Plex, Trakt, and Overseerr watchlists and automatically fetches requested movies and shows from torrent … | 10 | 1644 | active |
| Jesseovo/last30days-skill-cn An AI Agent skill (for Claude Code / OpenClaw) that automatically searches content from the last 30 days across 8 major Chinese internet pl… | 56 | 1642 | active |
| Sparticuz/chromium A TypeScript library that packages a serverless-optimized Chromium binary (Brotli-compressed) with decompression code and predefined launch… | 95 | 1641 | active |