function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| microsoft/Webwright Webwright is a Python framework from Microsoft that turns coding LLMs into state-of-the-art browser agents by giving them a terminal to lau… | 57 | 5947 | active |
| lanmaster53/recon-ng Recon-ng is a full-featured, modular reconnaissance framework for conducting web-based open source intelligence (OSINT) gathering. It offer… | 32 | 5871 | active |
| HapeLee/legado-with-MD3 Legado with MD3 is a free, open-source Android e-book and content aggregation reader, rebuilt from the Legado (阅读 3.0) project with a Mater… | 82 | 5860 | active |
| RedSiege/EyeWitness EyeWitness is a Python CLI tool that takes screenshots of websites using headless Chromium, captures server header information, and identif… | 50 | 5829 | active |
| joeseesun/qiaomu-anything-to-notebooklm A Claude Code Skill that ingests content from 15+ sources (WeChat articles, web pages, YouTube, PDFs, EPUB, Office docs, audio) and uploads… | 62 | 5824 | active |
| vantage-sh/ec2instances.info EC2Instances.info is an open-source web application for comparing Amazon EC2, RDS, and ElastiCache instance types, specs, and pricing acros… | 77 | 5759 | active |
| CharlesPikachu/musicdl A lightweight music downloader written in pure Python that supports dozens of music and audiobook platforms including NetEase Cloud Music, … | 75 | 5757 | active |
| omkarcloud/botasaurus Botasaurus is an all-in-one Python web scraping framework with built-in anti-detection, caching, parallelization, and proxy support. It let… | 72 | 5688 | active |
| vikiboss/60s A collection of free, open-source REST APIs delivering daily news digests, Chinese social media trending lists (Weibo, Bilibili, Douyin, Zh… | 92 | 5678 | active |
| qiye45/wechatVideoDownload A desktop GUI tool for downloading WeChat Channels (视频号) videos, live streams, live replays, and images. It automatically monitors WeChat t… | 71 | 5667 | active |
| TechXueXi TechXueXi is an open-source Python automation tool that automatically completes daily study tasks and quizzes on China's Xuexi Qiangguo (学习… | 56 | 5651 | active |
| rmax/scrapy-redis Redis-based components for Scrapy that enable distributed crawling and scraping by sharing a Redis queue across multiple spider instances. … | 60 | 5642 | active |
| executeautomation/mcp-playwright A Model Context Protocol (MCP) server that exposes Playwright browser automation to LLM clients like Claude Desktop, Cline, and Cursor IDE.… | 47 | 5635 | active |
| gosom/google-maps-scraper An open-source Go tool that scrapes Google Maps to extract business data such as names, addresses, phone numbers, websites, ratings, review… | 92 | 5630 | active |
| sallar/github-contributions-chart A web application that generates a shareable image of a user's entire GitHub contribution history, with multiple color themes. It works aro… | 35 | 5604 | active |
| potree/potree Potree is a free open-source WebGL-based renderer for viewing large point clouds directly in web browsers. It uses octree-based level-of-de… | 50 | 5585 | active |
| bilawalsidhu/gods-eye-view God's Eye View is a browser-based photorealistic 3D globe application that visualizes live public data feeds — aircraft, ships, satellites,… | 58 | 5574 | active |
| xuejianxianzun/PixivBatchDownloader A browser extension (Chrome, Edge, Firefox) for batch downloading illustrations, manga, Ugoira animations, and novels from Pixiv. It offers… | 98 | 5564 | active |
| AngleSharp/AngleSharp AngleSharp is a .NET library that parses HTML5, SVG, MathML, XML, and CSS into a fully standards-compliant W3C DOM. It supports querySelect… | 99 | 5531 | active |
| tangyoha/telegram_media_downloader A cross-platform Python application that downloads media files (video, audio, photos, documents) from Telegram chats, channels, and private… | 53 | 5495 | active |
| tebelorg/RPA-Python A Python package for robotic process automation (RPA) that wraps TagUI to automate web pages, desktop apps, and visual elements via a simpl… | 65 | 5493 | active |
| browser-act/skills BrowserAct Skills is a Python-based browser automation CLI designed for AI agents, providing real-browser control with anti-bot evasion (st… | 60 | 5446 | active |
| jef/streetmerchant streetmerchant is a Node.js application that continuously checks retail websites for product stock availability, with optional add-to-cart … | 64 | 5402 | active |
| Neet-Nestor/Telegram-Media-Downloader A browser userscript (installable via Tampermonkey, Violentmonkey, etc.) that adds media download buttons to the Telegram webapp. It enable… | 64 | 5253 | active |
| AhmadIbrahiim/Website-downloader A Node.js web application that downloads the complete source code of any website, including all assets like JavaScripts, stylesheets, and i… | 76 | 5245 | active |
| spatie/browsershot A PHP library that converts web pages or raw HTML into images, PDFs, or rendered HTML strings using Puppeteer and headless Chrome. It suppo… | 91 | 5241 | active |
| Yuukiy/JavSP JavSP is a Python command-line tool that scrapes adult video (JAV) metadata from multiple websites, aggregates the data, and generates NFO … | 25 | 5136 | active |
| scinfu/SwiftSoup SwiftSoup is a pure Swift HTML parser library that conforms to the WHATWG HTML5 specification and offers DOM traversal, CSS selectors, and … | 99 | 5119 | stable |
| hakluke/hakrawler Hakrawler is a fast command-line web crawler written in Go, built on the Gocolly library, that discovers URLs and JavaScript file locations… | 66 | 5116 | active |
| obsidianmd/obsidian-clipper Obsidian Web Clipper is the official browser extension for Obsidian that lets users highlight web pages and capture content as durable Mark… | 83 | 5092 | active |
| lc/gau gau (getallurls) is a Go CLI tool that fetches known URLs for a given domain from AlienVault's Open Threat Exchange, the Wayback Machine, C… | 56 | 5076 | active |
| apify/apify-mcp-server The Apify MCP Server exposes thousands of Apify Store scrapers, crawlers, and automation tools to AI agents via the Model Context Protocol,… | 84 | 5059 | active |
| binbyu/Reader A lightweight, free, open-source Windows (win32) reader application for txt and epub files, with support for reading online novels via conf… | 65 | 5057 | active |
| FellouAI/eko Eko is a production-ready JavaScript framework for building reliable AI agents and agentic workflows from natural language, running in both… | 63 | 4951 | active |
| FxEmbed/FxEmbed FxEmbed is a Cloudflare Worker service that powers FxTwitter, FixupX, and FxBluesky, rewriting X/Twitter and Bluesky links into rich embeds… | 77 | 4948 | active |
| exa-labs/exa-mcp-server An open-source MCP (Model Context Protocol) server that connects AI assistants and agents to Exa's web search, code search, content fetchin… | 66 | 4926 | active |
| jaypyles/Scraperr Scraperr is a self-hosted web scraping application with a web UI that lets users scrape websites without writing code, using XPath-based ex… | 10 | 4910 | active |
| MechanicalSoup/MechanicalSoup A Python library for automating interaction with websites, built on Requests and BeautifulSoup. It handles cookies, redirects, link followi… | 66 | 4888 | active |
| UndeadSec/SocialFish SocialFish is a Python-based phishing toolkit that clones modern login pages using Playwright browser automation and captures credentials, … | 76 | 4851 | active |
| fb55/htmlparser2 htmlparser2 is a fast, forgiving HTML and XML parser for JavaScript and TypeScript, offering a low-allocation callback interface as well as… | 87 | 4788 | stable |
| davidarroyo1234/InstagramUnfollowers A browser-based script that scans your Instagram account to identify users who don't follow you back, letting you selectively unfollow them… | 75 | 4781 | active |
| bjesus/pipet Pipet is a command-line web scraper written in Go that extracts data from online assets using HTML parsing, JSON parsing, and client-side J… | 24 | 4772 | active |
| Integuru-AI/Integuru Integuru is an AI agent that reverse-engineers platforms' internal APIs by analyzing browser network requests (HAR files) and building depe… | 62 | 4760 | active |
| l0o0/translators_CN A community-maintained collection of Zotero translators for Chinese academic and general websites, enabling Zotero to scrape citation metad… | 76 | 4721 | active |
| 88lin/video_vip A Tampermonkey/Greasemonkey userscript that integrates multiple third-party parsing interfaces to bypass VIP membership restrictions on Chi… | 63 | 4718 | active |
| DedSecInside/TorBot TorBot is a Python CLI tool for OSINT on the dark web, crawling .onion sites over the Tor network and building link trees. It can save craw… | 97 | 4716 | active |
| 201206030/novel-plus novel-plus is a full-featured novel/fiction CMS built on Spring Boot, comprising a reader-facing portal, author backend, admin dashboard, a… | 88 | 4709 | active |
| xroche/httrack HTTrack is a free offline browser utility that recursively downloads websites to a local directory, rewriting links so the mirrored copy ca… | 99 | 4702 | stable |
| ultrafunkamsterdam/nodriver Nodriver is a fully asynchronous Python browser automation and web scraping library, and the official successor to Undetected-Chromedriver.… | 62 | 4699 | active |
| lecepin/WeChatVideoDownloader A convenient desktop GUI application for downloading videos from WeChat Channels (WeChat Video Accounts). It intercepts and captures video … | 10 | 4677 | active |
| danburzo/percollate Percollate is a Node.js command-line tool that converts web pages into readable, well-formatted PDF, EPUB, HTML, or Markdown documents. It … | 53 | 4667 | active |
| ShadowHackrs/gmail-account-creator A Python-based automation tool for bulk creation of Gmail accounts, featuring anti-detection techniques like human-like typing simulation a… | 56 | 4657 | active |
| d60/twikit Twikit is a free Python library that wraps Twitter's internal API, allowing posting, searching, and scraping tweets without an official API… | 60 | 4631 | active |
| dataabc/weibo-crawler A Python crawler for Sina Weibo that scrapes user profiles and posts, exporting data to CSV, JSON, MySQL, MongoDB, or SQLite, and optionall… | 74 | 4625 | active |
| Keiyoushi Extensions A community-maintained repository of extensions (APKs) for Mihon and its forks, providing manga source plugins. The source code for the ext… | 71 | 4595 | active |
| linsomniac/spotify_to_ytmusic A set of Python scripts and a GUI for copying liked songs and playlists from Spotify to YouTube Music. It uses the Spotify backup data and … | 63 | 4582 | active |
| syhyz1990/baiduyun A free open-source Tampermonkey userscript that extracts real direct download links from Chinese cloud storage services (Baidu Netdisk, Ali… | 32 | 4505 | active |
| limbopro/Adblock4limbo A userscript-based ad-blocking project that removes popups, banners, and video ads on specific streaming, comic, novel, and adult sites via… | 77 | 4491 | active |
| sensepost/gowitness gowitness is a Go-based command-line utility that uses Chrome Headless to take screenshots of websites, supporting scans of URL lists, CIDR… | 75 | 4488 | active |
| 6dylan6/jdpro A collection of JavaScript automation scripts for the Qinglong panel, primarily for running scheduled JD (Jingdong) sign-in and reward task… | 73 | 4482 | active |
| metatube-community/jellyfin-plugin-metatube A metadata provider plugin for Jellyfin and Emby media servers that fetches movie and actor metadata from various internet providers via th… | 61 | 4454 | active |
| joeyism/linkedin_scraper A Python library that scrapes LinkedIn for user, company, and job data using Playwright with an async API. It provides Pydantic data models… | 82 | 4452 | active |
| sparklemotion/mechanize Mechanize is a Ruby library for automating interaction with websites. It handles cookies, redirects, link following, and form submission wh… | 91 | 4439 | stable |
| GerbenJavado/LinkFinder LinkFinder is a Python CLI script that discovers endpoints and their parameters in JavaScript files using jsbeautifier and regular expressi… | 32 | 4439 | stable |
| rachelos/we-mp-rss A self-hosted WeChat official account (公众号) subscription assistant that scrapes articles, generates RSS feeds, and converts content to Mark… | 84 | 4396 | active |
| VonChange/utao Utao TV (油桃TV, now upgraded to 土拨鼠浏览器) is a third-party browser designed for Android TV boxes and smart TVs that lets users watch live CCTV… | 67 | 4340 | active |
| QasimWani/LeetHub LeetHub is a browser extension that automatically pushes your LeetCode solutions to GitHub whenever you pass all tests on a problem. It sup… | 72 | 4339 | active |
| anasty17/mirror-leech-telegram-bot A Python Telegram bot that mirrors or leeches files from direct links, torrents, NZB/Usenet, Google Drive, rclone clouds, and yt-dlp/JDownl… | 77 | 4279 | active |
| UltimaHoarder/UltimaScraper A Python-based scraper that downloads all media (photos, videos) from OnlyFans accounts using the user's own session authentication. It sto… | 23 | 4271 | active |
| wasi-master/13ft A self-hosted web service that bypasses paywalls on news sites by fetching pages as GoogleBot, a replacement for the defunct 12ft.io. It se… | 83 | 4266 | active |
| hoothin/UserScripts A collection of Greasemonkey/Tampermonkey userscripts by hoothin, including Pagetual (auto-pager infinite scrolling), Picviewer CE+ (online… | 76 | 4264 | active |
| codeceptjs/CodeceptJS CodeceptJS is a Node.js end-to-end testing framework with a BDD-style syntax where tests are written as user-perspective scenarios using an… | 98 | 4241 | active |
| megadose/toutatis Toutatis is a Python CLI tool that extracts public information from Instagram accounts, such as emails, phone numbers, follower counts, and… | 32 | 4240 | active |
| praw-dev/praw PRAW is a Python package that wraps Reddit's API, providing simple, rule-compliant access to Reddit data and actions. It handles OAuth auth… | 98 | 4234 | stable |
| morpheus65535/bazarr Bazarr is a self-hosted companion application to Sonarr and Radarr that automatically manages and downloads subtitles for your TV series an… | 92 | 4234 | active |
| watsonbox/exportify Exportify is a browser-based application for exporting and backing up Spotify playlists to CSV files using the Spotify Web API. It runs ent… | 65 | 4201 | active |
| vogler/free-games-claimer A Node.js automation tool that periodically claims free games and DLCs on the Epic Games Store, Amazon Prime Gaming, and GOG, plus free Unr… | 73 | 4197 | active |
| kanasimi/work_crawler A multi-language downloader application that batch-downloads web novels (converting them to EPUB) and comics from a large list of Chinese, … | 66 | 4193 | active |
| Patchright Patchright is a patched, undetected fork of the Playwright browser automation framework that evades bot-detection systems like Cloudflare. … | 93 | 4191 | active |
| InstaPy/InstaPy InstaPy is a Python library built on Selenium that automates Instagram interactions such as liking, commenting, following, and unfollowing … | 27 | 18167 | maintenance |
| speedyapply/JobSpy JobSpy is a Python library that scrapes job postings from popular job boards like LinkedIn, Indeed, Glassdoor, Google, and ZipRecruiter con… | 57 | 4165 | active |
| 027xiguapi/code-box CodeBox is a browser extension for Chrome, Edge, Firefox, and 360 browsers that enhances Chinese tech blog sites like CSDN, Zhihu, Juejin, … | 60 | 4140 | active |
| dotnetcore/DotnetSpider DotnetSpider is a .NET Standard web crawling and scraping framework that is lightweight, efficient, and cross-platform. It supports distrib… | 57 | 4138 | active |
| ivre/ivre IVRE is an open-source network recon framework written in Python that collects, stores, and analyzes network intelligence from active scann… | 66 | 4119 | active |
| ericciarla/trendFinder A self-hosted Node.js application that monitors influencer posts on Twitter/X and website changes via Firecrawl, then uses LLMs (Together A… | 24 | 4116 | active |
| Lucksi/Mr.Holmes Mr.Holmes is a Python-based OSINT (open-source intelligence) CLI tool that gathers information about usernames, domains, phone numbers, and… | 53 | 4112 | active |
| nghuyong/WeiboSpider A continuously maintained Python web scraping tool for Sina Weibo built on Scrapy and the new weibo.com API. It collects user profiles, pos… | 73 | 4109 | active |
| RipMeApp/ripme RipMe is a cross-platform Java application that bulk-downloads image albums from websites like Reddit, Imgur, Twitter, Instagram, and Tumbl… | 98 | 4104 | active |
| jasonxtn/Argus Argus is a Python-based all-in-one information gathering and reconnaissance toolkit with an interactive console and modular architecture. I… | 47 | 4082 | active |
| IonicaBizau/scrape-it scrape-it is a Node.js web scraping library with a simple, declarative API for extracting data from HTML pages, built on top of tinyreq and… | 93 | 4073 | active |
| ipcjs/oh-my-userscripts A collection of Tampermonkey/Greasyfork userscripts by ipcjs that tweak and enhance websites like Bilibili, Zhihu, Bangumi, S1, and Google.… | 73 | 4065 | active |
| 5rahim/seanime Seanime is an open-source self-hosted media server for anime and manga, offering a web interface and desktop app to manage local libraries,… | 86 | 4062 | active |
| React Native Upgrade Helper A web application that shows the exact file diffs between any two React Native versions to guide app upgrades. It is built on the rn-diff-p… | 75 | 4062 | active |
| fake-useragent/fake-useragent A Python library that generates realistic, up-to-date browser user-agent strings from a bundled real-world database. It supports random or … | 10 | 4049 | active |
| Guyungy/damaihelper DamaiHelper is a multi-platform ticket-grabbing automation assistant (Damai, Taopiaopiao, Binwandao) built as a Python backend with an Ant … | 74 | 4042 | active |
| hafrey1/LunaTV-config A configuration repository and Cloudflare Workers-based CORS proxy for MoonTV/LunaTV video source APIs, with daily automated API health che… | 63 | 4040 | active |
| nkanaev/yarr yarr is a web-based RSS feed aggregator written in Go, distributed as a single dependency-free binary that runs as a desktop app or self-ho… | 85 | 4031 | active |
| symfony/dom-crawler Symfony DomCrawler is a PHP component that eases DOM navigation for HTML and XML documents. It provides a Crawler class for querying and tr… | 99 | 4027 | stable |
| libredirect/browser_extension LibRedirect is a browser extension (WebExtension) that automatically redirects requests to popular sites like YouTube, Twitter, Reddit, Ins… | 88 | 4024 | active |
| imsyy/DailyHotApi DailyHotApi is a TypeScript-based API service that aggregates trending/hot list data from many Chinese platforms (Bilibili, Weibo, Zhihu, D… | 61 | 4021 | active |