function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| GramAddict/bot GramAddict is a free, open-source Instagram bot that automates liking, following, commenting, and messaging on any Android device or emulat… | 26 | 1632 | active |
| vigorX777/ai-daily-digest A zero-dependency TypeScript CLI tool (runnable as an OpenCode skill) that fetches the latest posts from 90 top tech blogs recommended by A… | 45 | 1630 | active |
| lit26/finvizfinance A Python library that scrapes and downloads financial data from the FinViz website, returning stock fundamentals, technicals, charts, news,… | 79 | 1623 | active |
| vladko312/SSTImap SSTImap is a Python-based penetration testing tool that automatically detects and exploits Server-Side Template Injection (SSTI) and code i… | 88 | 1621 | active |
| michenriksen/aquatone Aquatone is a Go CLI tool for visual inspection of websites across many hosts, taking screenshots via headless Chrome/Chromium and generati… | 10 | 5961 | maintenance |
| wu529778790/panhub.shenzjd.com PanHub is a self-hostable netdisk search aggregator that combines results from Quark, Aliyun Drive, Baidu Netdisk, 115, Thunder and 80+ Tel… | 62 | 1617 | active |
| hartator/wayback-machine-downloader A Ruby command-line tool that downloads an entire website from the Internet Archive Wayback Machine, restoring original files and directory… | 23 | 5930 | maintenance |
| qsniyg/maxurl Image Max URL is a userscript and browser extension that finds larger or original versions of images and videos by rewriting URL patterns, … | 77 | 1611 | active |
| ka-pi-ba-la/AIbijia AIbijia is a price-comparison website and community resource that aggregates AI subscription and token prices across multiple platforms (Ch… | 57 | 1611 | active |
| minkphp/Mink Mink is a PHP library providing an abstraction layer over web browser emulators and drivers (like Goutte, Selenium) for controlling browser… | 53 | 1611 | active |
| ArchiveTeam/grab-site grab-site is a preconfigured web crawler for archiving websites, producing WARC files via a fork of wpull. It includes a dashboard for moni… | 43 | 1607 | active |
| matthewmueller/x-ray x-ray is a Node.js web scraping library that lets you define flexible schemas to structure data from any website using jQuery-like selector… | 66 | 5907 | maintenance |
| Altimis/Scweet Scweet is a Python library and CLI for scraping tweets, profile timelines, followers, following lists, and user profiles from Twitter/X wit… | 78 | 1604 | active |
| Xatta-Trone/medium-parser-extension A browser extension for Chrome, Edge, and Firefox that lets users read member-only articles on medium.com and Medium-based sites (e.g. Towa… | 59 | 1589 | active |
| lanyeeee/jmcomic-downloader A multi-threaded GUI downloader for the 18comic.vip (jmcomic) manga site, built with Tauri (Rust backend, Vue frontend). It supports search… | 89 | 1587 | active |
| Autumn-27/ScopeSentry ScopeSentry is a self-hosted attack surface and asset mapping platform that combines subdomain enumeration, port scanning, fingerprinting, … | 88 | 1587 | active |
| s0md3v/uro uro is a Python CLI tool that declutters URL lists for crawling and security testing without making any HTTP requests. It removes duplicate… | 29 | 1587 | stable |
| xnl-h4ck3r/xnLinkFinder xnLinkFinder is a Python CLI tool that discovers endpoints, potential parameters, target-specific wordlists, and secrets for a given target… | 78 | 1585 | active |
| m3n0sd0n4ld/GooFuzz GooFuzz is a Bash-based CLI tool that performs fuzzing-style reconnaissance using advanced Google searches (Google Dorking) via the Google … | 57 | 1585 | active |
| LeyckerS/moondownloader Moon Downloader is a bulk file downloader for the datanodes.to and fuckingfast.co file hosts, extracting direct links either via a real Chr… | 88 | 1584 | active |
| actionbook/actionbook Actionbook is an AI agent that lives in a Chrome side panel and operates websites on your behalf, reading data behind logins and paywalls a… | 74 | 1584 | active |
| AlisamTechnology/ATSCAN ATSCAN is a Perl-based command-line scanner for mass dork searching and vulnerability exploitation. It combines search engine dorking with … | 23 | 1583 | active |
| m8sec/CrossLinked CrossLinked is a Python CLI tool that enumerates LinkedIn employee names for an organization by scraping search engine results, without nee… | 23 | 1582 | active |
| agentrhq/webcmd Webcmd is a self-learning browser infrastructure CLI for AI agents that compiles knowledge of websites into deterministic commands, cutting… | 80 | 1579 | active |
| postlight/parser Postlight Parser is a JavaScript library that extracts meaningful content—article text, titles, authors, dates, lead images, and excerpts—f… | 23 | 5786 | maintenance |
| Leon406/SubCrawler A Kotlin-based tool that automatically crawls and health-checks (via Google ping) free public proxy nodes for protocols like V2Ray, Shadows… | 77 | 1561 | active |
| OWASP/QRLJacking QRLJacking is an OWASP project documenting and exploiting the Quick Response Code Login Jacking attack vector, which hijacks user sessions … | 48 | 1559 | active |
| ion-design/ditto.site ditto.site is a hosted capture-to-code service that turns a public URL into a self-contained TypeScript app, emitting a deterministic Next.… | 57 | 1558 | active |
| ttttmr/Wechat2RSS Wechat2RSS is a service and self-hostable tool that converts WeChat official account (公众号) articles into RSS feeds, aiming for updates with… | 75 | 1557 | active |
| tafia/quick-xml quick-xml is a high-performance XML pull reader and writer library for Rust, featuring near zero-copy parsing with Cow types and buffer reu… | 95 | 1556 | stable |
| webrecorder/archiveweb.page ArchiveWeb.page is a high-fidelity web archiving tool that runs as a Chrome/Chromium browser extension and as a standalone Electron app. It… | 93 | 1555 | active |
| ulixee/hero Hero is a headless web browser built specifically for web scraping, powered by Chrome and controlled from NodeJS with a fully compliant DOM… | 69 | 1554 | active |
| littledivy/mimic mimic is a Python tool that captures traffic from mobile or web apps via mitmproxy, extracts authentication material, and uses AI to genera… | 55 | 1550 | active |
| Rhizome-Conifer/conifer Conifer is an open-source web archiving platform for capturing, replaying, and sharing collections of archived web pages through a user-fri… | 74 | 1549 | active |
| skernelx/tavily-key-generator A Python toolkit that automates signup flows for Tavily, Firecrawl, and Exa using real browser automation (Playwright/Camoufox), Turnstile … | 48 | 1549 | active |
| yujiosaka/headless-chrome-crawler A Node.js library providing a distributed web crawler powered by Headless Chrome via Puppeteer. It can crawl JavaScript-rendered (SPA) webs… | 23 | 5635 | maintenance |
| ecmadao/hacknical Hacknical is a web application that analyzes a GitHub user's data (contributions, commits, languages, repos) and helps generate a better de… | 49 | 1542 | active |
| hyperbrowserai/HyperAgent HyperAgent is a TypeScript library and CLI that adds LLM-powered natural language commands to Playwright for browser automation. It support… | 52 | 1540 | active |
| DialmasterOrg/Youtarr Youtarr is a self-hosted web application that automatically downloads YouTube channel and playlist content, organizes it with metadata for … | 92 | 1539 | active |
| oxylabs/ai-map-py AI-Map is a Python SDK client for Oxylabs AI Studio's AI-powered website mapping service, which discovers and extracts relevant URLs from a… | 50 | 1537 | active |
| FinanceData/FinanceDataReader FinanceDataReader is a Python library and CLI tool for reading financial data such as stock listings, stock prices, indexes, exchange rates… | 69 | 1535 | active |
| cross-seed/cross-seed cross-seed is a Node.js application that automatically finds and downloads torrents matching your existing torrent library across multiple … | 92 | 1534 | active |
| LifeActor/ykdl YouKuDownLoader (ykdl) is a Python command-line video downloader focused on China mainland video sites, forked from you-get with restructur… | 47 | 1533 | active |
| xnl-h4ck3r/GAP-Burp-Extension GAP is a Burp Suite extension written in Python (Jython) that extracts potential endpoints, parameters, and links from Burp's site map, pro… | 66 | 1530 | active |
| rbren/rss-parser rss-parser is a lightweight JavaScript library that converts RSS and Atom XML feeds into JavaScript objects. It works in both Node.js and t… | 77 | 1527 | active |
| justfoolingaround/animdl animdl is a lightweight Python CLI tool that scrapes, streams, and downloads anime episodes from supported providers. It supports quality s… | 32 | 1522 | active |
| tidyverse/rvest rvest is an R package from the tidyverse for scraping (harvesting) data from web pages, inspired by Beautiful Soup and RoboBrowser. It prov… | 43 | 1520 | active |
| SpiderClub/haipproxy A high-availability distributed IP proxy pool built with Scrapy and Redis that scrapes free proxies from the internet, validates them, and … | 23 | 5523 | maintenance |
| jasperan/whatsapp-osint A Python CLI tool that uses Selenium to track when WhatsApp contacts go online/offline, logging presence sessions to SQLite. It exports dat… | 76 | 1511 | active |
| Pickle-Pixel/ApplyPilot ApplyPilot is an open-source AI agent that autonomously applies to jobs on your behalf across any job site or application form. It runs a 6… | 60 | 1508 | active |
| oxylabs/browser-agent-py A Python SDK for Oxylabs AI Studio's Browser Agent, a cloud service that automates real-user browsing tasks (clicking, typing, scrolling, s… | 51 | 1505 | active |
| manojVivek/medium-unlimited A browser extension for Chrome and Firefox that unlocks Medium.com membership-only articles by bypassing the paywall. It supports medium.co… | 10 | 5467 | maintenance |
| JustAnotherArchivist/snscrape snscrape is a Python-based scraper for social networking services that extracts posts, profiles, hashtags, and search results from platform… | 32 | 5444 | maintenance |
| rafatosta/zapzap ZapZap is an unofficial WhatsApp Web desktop client built with Python, PyQt6, and QtWebEngine that wraps web.whatsapp.com in a native deskt… | 95 | 1495 | active |
| Vincentqyw/cv-arxiv-daily An automated daily digest of computer vision and robotics arXiv papers (SLAM, SFM, visual localization, keypoint detection, image matching,… | 77 | 1494 | active |
| SilentDemonSD/WZML-X WZML-X is a self-hosted Telegram mirror and leech bot written in Python that downloads files from torrents, Mega, Google Drive, direct link… | 96 | 1493 | active |
| jlesage/docker-jdownloader-2 A Docker container packaging JDownloader 2, a download manager, with a browser-accessible graphical interface via web or VNC. It is an unof… | 98 | 1492 | active |
| Tsuk1ko/bilibili-live-chat A backend-free web app that displays Bilibili live stream danmaku (chat messages) and gifts in a YouTube Live Chat style overlay, primarily… | 60 | 1486 | active |
| exa-labs/company-researcher An open-source web application by Exa.ai that lets users enter a company URL and instantly gathers comprehensive research about it, includi… | 65 | 1483 | active |
| AI4Finance-Foundation/FinNLP FinNLP is a Python library for collecting internet-scale financial data from sources like Finnhub, Yahoo Finance, Reuters, and Sina Finance… | 31 | 1480 | active |
| sethblack/python-seo-analyzer A Python-based SEO analyzer that crawls a website, analyzes its structure, counts body words, and reports technical SEO issues, with option… | 65 | 1475 | active |
| Ademking/MD-This-Page A browser extension for Chrome and Firefox that converts any webpage into clean, readable Markdown with one click, using Mozilla's Readabil… | 61 | 1475 | active |
| Surfer-Org/Protocol Surfer Protocol is an open-source framework for exporting personal data from platforms like Gmail, iMessages, Twitter, Notion, and ChatGPT.… | 13 | 1475 | active |
| shuanx/BurpAPIFinder BurpAPIFinder is a Burp Suite extension written in Java that passively analyzes HTTP traffic (HTML and JS files) to discover hidden API end… | 15 | 1472 | active |
| epsylon/xsser XSSer is an automatic penetration testing framework for detecting, exploiting, and reporting cross-site scripting (XSS) vulnerabilities in … | 82 | 1461 | active |
| momosecurity/FindSomething FindSomething is a passive browser extension for Chrome and Firefox that extracts potentially sensitive information (like emails, API keys,… | 32 | 1458 | active |
| mylar3/mylar3 Mylar3 is a Python-based automated comic book (cbr/cbz) downloader that monitors a watchlist of series and grabs new issues via NZB indexer… | 64 | 1455 | active |
| roach-php/core Roach is a complete web scraping and crawling toolkit for PHP, heavily inspired by Python's Scrapy. It lets developers define spiders that … | 45 | 1455 | active |
| tychxn/jd-assistant A JD.com (Jingdong) purchase assistant written in Python that automates login via QR code, product stock/price queries, cart management, an… | 32 | 5259 | maintenance |
| submato/xhscrawl A Python-based reverse-engineering toolkit for Xiaohongshu (XHS) web APIs, focusing on generating the encrypted x-s signature parameter via… | 72 | 1452 | active |
| tinyfish-io/agentql AgentQL is a suite of tools for extracting structured data and automating workflows on live websites using an AI-powered natural language q… | 70 | 1451 | active |
| gpodder/gpodder gPodder is a free, open-source podcast client and media aggregator that lets users subscribe to, download, and manage podcast feeds. It has… | 66 | 1451 | stable |
| eliasdabbas/advertools advertools is a Python library of productivity and analysis tools for online/digital marketing, built as a set of independent, composable f… | 73 | 1446 | active |
| orangecoding/fredy Fredy is a self-hosted Node.js application that continuously scrapes European real estate portals like ImmoScout24, Immowelt, Kleinanzeigen… | 95 | 1444 | active |
| 6551Team/opentwitter-mcp A Python MCP server that exposes Twitter/X data (user profiles, tweet search, follower events, deleted tweets, KOL tracking) to AI assistan… | 59 | 1444 | active |
| denandz/sourcemapper Sourcemapper is a Go CLI tool that parses JavaScript sourcemap (.map) files, whether from URLs or local directories, and reconstructs the o… | 75 | 1442 | stable |
| txperl/PixivBiu PixivBiu is a self-hosted Pixiv client written in Go that provides a browser-based interface for searching, filtering, browsing, and downlo… | 87 | 1440 | active |
| scrapy-plugins/scrapy-playwright A Scrapy download handler that uses Playwright for Python to fetch pages, enabling scraping of JavaScript-rendered sites while keeping the … | 86 | 1439 | active |
| nickclyde/duckduckgo-mcp-server A Model Context Protocol (MCP) server that exposes DuckDuckGo web search to LLM clients like Claude Desktop and Claude Code. It also fetche… | 83 | 1439 | active |
| asz798838958/aBaiFreeGPT A self-hosted account registration and lifecycle management platform that automates bulk account signup, email verification, TOTP 2FA bindi… | 58 | 1438 | active |
| ShilongLee/Crawler A self-hostable crawler API server that exposes HTTP endpoints for scraping public data from Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo… | 53 | 1435 | active |
| apinanaivot/IKEA-3D-Model-Download-Button A Tampermonkey userscript that adds a 'Download 3D' button to IKEA product pages, letting users save the product's 3D model as a .GLB file.… | 53 | 1435 | active |
| drawrowfly/tiktok-scraper A TypeScript library and CLI tool that scrapes TikTok metadata from user, hashtag, trend, and music pages and downloads video posts without… | 23 | 5173 | maintenance |
| sw33tLie/bbscope bbscope is a Go CLI tool that fetches, stores, and manages bug bounty program scopes from HackerOne, Bugcrowd, Intigriti, YesWeHack, and Im… | 73 | 1432 | active |
| wd210010/only_for_happly A collection of Python automation scripts designed to run on the Qinglong (青龙) panel, covering daily check-ins for services like Baidu Tieb… | 74 | 1431 | active |
| dteviot/WebToEpub A Chrome and Firefox browser extension that converts web novels and other web pages into EPUB ebooks for offline reading. It supports many … | 94 | 1430 | active |
| kepano/clipper-templates A collection of templates for the Obsidian Web Clipper browser extension, covering generic schemas (recipes, products) and specific sites l… | 55 | 1427 | active |
| zhaoolee/garss Garss (嘎!RSS) is a self-hosted RSS aggregation and reading system that uses GitHub Actions to collect hundreds of RSS feeds and render them… | 82 | 1426 | active |
| lanlinju/Animius Animius is a clean, minimalist Android app for watching anime, built with Jetpack Compose and Kotlin. It supports danmaku (bullet comments)… | 82 | 1425 | active |
| rebrowser/rebrowser-patches A collection of source-code patches for Puppeteer and Playwright that fix automation leaks and help avoid bot detection systems like Cloudf… | 34 | 1424 | active |
| stereobooster/react-snap A zero-configuration, framework-agnostic static prerendering tool for single-page applications. It uses Headless Chrome (via Puppeteer) to … | 62 | 5116 | maintenance |
| lorenzodifuccia/safaribooks A Python CLI tool that downloads books from O'Reilly Learning (Safari Books Online) and generates EPUB files from them. It requires a valid… | 51 | 5100 | maintenance |
| openwpm/OpenWPM OpenWPM is a web privacy measurement framework built on Firefox with Selenium automation, designed to collect data from thousands to millio… | 95 | 1417 | active |
| danny0838/content-farm-terminator Content Farm Terminator is a cross-platform browser extension that identifies content farms by marking hyperlinks pointing to them and bloc… | 77 | 1416 | active |
| debridmediamanager/debrid-media-manager Debrid Media Manager (DMM) is a free, open-source web app for curating an unlimited-size movie and TV show library on debrid services like … | 75 | 1416 | active |
| hxh19950701/WebViewTvLive An Android TV live streaming app that loads official broadcaster web pages in a Tencent X5 WebView, auto-fullscreens the video element, and… | 87 | 5094 | maintenance |
| facundoolano/app-store-scraper A Node.js library for scraping application data from the iTunes and Mac App Store. It provides methods to retrieve app details, search resu… | 47 | 1414 | active |
| sec-edgar/sec-edgar A Python library and CLI for downloading company periodic reports, filings, and forms from the SEC's EDGAR database. It supports fetching f… | 48 | 1413 | active |
| cdpdriver/zendriver Zendriver is an async-first Python web scraping and browser automation framework built on the Chrome Devtools Protocol, forked from nodrive… | 88 | 1409 | active |
| browserwing/browserwing BrowserWing is an open-source browser automation platform written in Go with a React dashboard that exposes browser control via MCP command… | 71 | 1408 | active |