function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| blacklanternsecurity/MANSPIDER MANSPIDER is a Python CLI tool that crawls SMB shares across entire networks to find files by filename or content, with regex support and t… | 77 | 1406 | active |
| superhedgy/AttackSurfaceMapper AttackSurfaceMapper is a Python CLI reconnaissance tool that expands a target's attack surface using OSINT and active techniques like subdo… | 32 | 1405 | active |
| nianzhibai/91 A self-hosted private video site written in Go that aggregates videos from multiple cloud drives (115, PikPak, 123pan, OneDrive, Google Dri… | 80 | 1404 | active |
| KoalaBear84/OpenDirectoryDownloader A cross-platform C#/.NET command-line tool that indexes open directory listings across 130+ supported formats, including FTP(S), Google Dri… | 94 | 1389 | active |
| lorey/mlscraper mlscraper is a Python library that automatically extracts structured data from HTML pages using machine learning. Instead of writing CSS se… | 23 | 1385 | active |
| Avnsx/fansly-downloader A Python-based tool for bulk downloading photos, videos, and audio from fansly.com, also shipped as a standalone Windows executable. It sup… | 10 | 1384 | active |
| thinh-vu/vnstock Vnstock is an open-source Python library for extracting and analyzing Vietnam stock market data, returning data as pandas DataFrames via si… | 90 | 1383 | active |
| jachinlin/geektime_dl A Python CLI tool that downloads Geektime (极客时间) courses and converts them into ebooks for reading on Kindle. It handles login, course quer… | 66 | 1381 | active |
| CIRCL/AIL-framework AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T… | 67 | 1378 | active |
| mattsse/chromiumoxide chromiumoxide is a Rust library providing a high-level async API for controlling Chrome or Chromium via the Chrome DevTools Protocol. It ca… | 75 | 1375 | active |
| okfn-brasil/querido-diario Querido Diário is an open-source project by Open Knowledge Brasil that scrapes and aggregates Brazilian municipal official gazettes (diário… | 75 | 1373 | active |
| tavily-ai/tavily-python The official Python SDK for the Tavily API, providing search, content extraction, crawling, site mapping, and research capabilities. It let… | 72 | 1373 | active |
| Minerchu/dongguaTV A self-hosted Node.js video aggregation platform that searches 30+ movie/TV resource site APIs, aggregates results, and scrapes metadata fr… | 40 | 1372 | active |
| xjbeta/iina-plus IINA+ is a small macOS application that adds danmaku (bullet comment) support and live-stream/video playback for Chinese platforms to the I… | 97 | 1370 | active |
| dethcrypto/dethcode DethCode lets you view the verified source of deployed Ethereum smart contracts in an ephemeral VS Code instance by changing an Etherscan U… | 52 | 1370 | active |
| fasnow/fine Fine is a Chinese-language cyberspace asset mapping and reconnaissance tool integrating FOFA, Hunter, Quake, ZoomEye, and Shodan APIs, plus… | 82 | 1366 | active |
| python273/vk_api vk_api is a Python library that wraps the VKontakte (vk.com) API, simplifying authentication and API method calls for building scripts and … | 83 | 1364 | active |
| xlang-ai/OpenAgents OpenAgents is an open platform for using and hosting LLM-powered language agents, featuring a Data Agent for Python/SQL analysis, a Plugins… | 28 | 4859 | maintenance |
| Maasea/sgmodule A collection of Surge modules (sgmodule files) written in JavaScript that enhance or modify mobile apps and network behavior, such as remov… | 74 | 1359 | active |
| xiaohucode/xiangse A curated collection of video and manga source plugins (.xbs files) for the Xiangse Guige (香色闺阁) reading/media app, imported via URL. It ag… | 10 | 1359 | active |
| SplashtopInc/winstall winstall is a free, open-source web app for browsing and searching Microsoft's Windows Package Manager (winget) repository of 14,600+ apps.… | 77 | 1358 | active |
| firecrawl/open-scouts Open Scouts is an AI-powered web monitoring platform where users create automated 'scouts' that run on a schedule to search the web and sen… | 53 | 1358 | active |
| fugary/calibre-douban A Calibre metadata source plugin that fetches book metadata from Douban by crawling book.douban.com web pages, since Douban no longer offer… | 63 | 1357 | active |
| scrapy/parsel Parsel is a BSD-licensed Python library for extracting data from HTML, XML, and JSON documents using CSS selectors, XPath expressions, JMES… | 88 | 1352 | active |
| raznem/parsera Parsera is a lightweight Python library for scraping websites using LLMs, letting users define elements to extract with natural-language de… | 54 | 1350 | active |
| nmdias/FeedKit FeedKit is a Swift library for parsing and generating RSS, Atom, and JSON Feed formats. It supports common namespaces like Dublin Core, Med… | 99 | 1348 | active |
| MiniGlome/Archive.org-Downloader A Python 3 command-line script that downloads borrowable books from archive.org and Open Library and assembles them into PDF files. It requ… | 76 | 1348 | active |
| LeetaoGoooo/RSSAid RSSAid is a Flutter-based mobile app that complements RSSHub by helping users discover and subscribe to RSS feeds from websites, similar to… | 76 | 1345 | active |
| philippta/flyscrape Flyscrape is a standalone command-line web scraping tool written in Go that lets users write extraction logic in JavaScript with a jQuery-l… | 42 | 1345 | active |
| SpiderClub/weibospider A distributed web crawler for Sina Weibo (Chinese microblogging platform) built with Python, Celery, and requests. It scrapes user profiles… | 32 | 4793 | maintenance |
| iota9star/mikan_flutter A third-party Flutter mobile client for the Mikan Project (mikanani.me), an anime/seasonal bangumi torrent subscription site. It lets users… | 99 | 1343 | active |
| mvdbos/php-spider A configurable and extensible PHP web spider library for crawling websites. It supports breadth-first and depth-first traversal, URI discov… | 75 | 1341 | active |
| jocmp/capyreader Capy Reader is a free, open-source RSS feed reader and news aggregator for Android built with Kotlin and Jetpack Compose. It syncs with Fee… | 89 | 1334 | active |
| DUpdateSystem/UpgradeAll UpgradeAll is a free, open-source Android app that checks for updates to installed Android apps, Magisk modules, and other software from a … | 67 | 1334 | active |
| taf2/curb Curb provides Ruby-language bindings for libcurl, the fully-featured client-side URL transfer library, supporting both easy and multi modes… | 75 | 1331 | active |
| mariostoev/finviz An unofficial Python library for scraping FinViz.com, providing stock data, news, insider transactions, analyst price targets, and a stock … | 60 | 1331 | active |
| ytdl node-ytdl-core is a JavaScript library for downloading YouTube videos in Node.js, exposing a stream-friendly API with format selection, byt… | 46 | 4730 | maintenance |
| tophubs/TopList TopList (今日热榜) is a self-hosted aggregation website that collects trending headlines from popular sites like Zhihu, Hupu, and V2EX. It is w… | 32 | 4730 | maintenance |
| sockysec/Telerecon Telerecon is a Python-based OSINT reconnaissance framework for researching and investigating Telegram. It scrapes user profiles, messages, … | 28 | 1324 | active |
| cinemagoer/cinemagoer Cinemagoer (formerly IMDbPY) is a Python package for retrieving and managing IMDb data about movies, people, characters, and companies from… | 95 | 1323 | active |
| MrTuxx/SocialPwned SocialPwned is a Python-based OSINT tool that harvests emails published on Instagram, LinkedIn, and Twitter to find credential leaks via Pw… | 10 | 1320 | active |
| monosans/proxy-scraper-checker A fast async Rust CLI tool that scrapes HTTP, SOCKS4, and SOCKS5 proxies from arbitrary text, HTML, or JSON sources, verifies each proxy by… | 77 | 1318 | active |
| lucahammer/tweetXer A JavaScript userscript that bulk-deletes all your tweets on X (Twitter) for free, using your official Data Export file. It runs in the bro… | 55 | 1316 | active |
| darwin-lau/langmanus LangManus is a community-driven Python framework for building hierarchical multi-agent AI automation systems, where a supervisor coordinate… | 25 | 1316 | active |
| Bloggify/github-calendar A JavaScript library that embeds a user's GitHub contributions calendar into any web page with a single function call. It fetches contribut… | 44 | 1315 | active |
| jmerle/competitive-companion A browser extension that parses competitive programming problems from online judges like Codeforces, AtCoder, and Kattis. It extracts test … | 87 | 1309 | active |
| codingo/VHostScan VHostScan is a Python-based virtual host scanner that discovers hidden vhosts on a web server using wordlists, reverse lookups, and catch-a… | 39 | 1309 | active |
| karust/openserp OpenSERP is a self-hosted, MIT-licensed SERP API and CLI written in Go that returns structured search results from Google, Bing, Yandex, Ba… | 91 | 1306 | active |
| yasserg/crawler4j crawler4j is an open-source web crawler library for Java that provides a simple interface for building multi-threaded web crawlers in minut… | 23 | 4618 | maintenance |
| dwisiswant0/go-dork go-dork is a fast command-line dork scanner written in Go that automates Google dorking across multiple search engines. It supports Google,… | 23 | 1301 | stable |
| DimiMikadze/orca Orca is an AI agent application for deep LinkedIn profile analysis that scrapes posts, comments, reactions, and interaction networks, then … | 75 | 1297 | active |
| bookstairs/bookhunter bookhunter is a Go command-line tool for scraping and downloading ebooks from sources like Talebook, SoBooks, Telegram channels, and China'… | 54 | 1295 | active |
| h4r5h1t/webcopilot WebCopilot is a Bash-based automation script for bug bounty reconnaissance that enumerates subdomains using multiple tools, filters paramet… | 23 | 1295 | active |
| yjl9903/AnimeGarden AnimeGarden is a third-party mirror and aggregation site for 動漫花園 (dmhy) anime BT torrents, offering a web UI, open REST API, RSS feeds, an… | 87 | 1294 | active |
| techwithtim/Price-Tracking-Web-Scraper A full-stack price tracking application that scrapes product prices (currently Amazon.ca) using Playwright and Bright Data's Scraping Brows… | 29 | 1294 | active |
| Steamauto/Steamauto Steamauto is a free, open-source Python application that fully automates buying, selling, and delivery of CS2/CSGO skins across Steam and C… | 96 | 1292 | active |
| foxhui/WebAI2API WebAI2API is a self-hosted Node.js service that exposes web-based AI services (LMArena, Gemini, ChatGPT, DeepSeek, etc.) as OpenAI-compatib… | 57 | 1292 | active |
| vega-org/vega-app Vega is an open-source Android media streaming app built with React Native and TypeScript that lets users stream and download video content… | 91 | 1290 | active |
| zohaibbashir/Google-Maps-Scrapper A Python CLI script built on Playwright that scrapes Google Maps listings to extract business details such as name, address, website, phone… | 65 | 1289 | active |
| oxylabs/paid-proxy-servers A promotional GitHub repository for Oxylabs' commercial paid proxy services, covering residential, mobile, datacenter, ISP, and SOCKS5 prox… | 59 | 1289 | active |
| P1-Team/AlliN AlliN is a flexible, dependency-free Python scanner designed to assist penetration testing projects, especially lateral movement and intran… | 40 | 1288 | active |
| misiektoja/instagram_monitor A Python-based OSINT tool that tracks Instagram users' activities in real time, including story updates, profile changes, and follower shif… | 91 | 1286 | active |
| denho/faved Faved is a free, open-source, self-hostable bookmark manager with customizable nested tags, instant search, and duplicate detection, built … | 83 | 1282 | active |
| dvcoolarun/web2pdf A Python command-line tool that converts webpages into formatted PDFs using WeasyPrint. It supports batch conversion, recursive same-domain… | 67 | 1281 | active |
| TheBeastLT/torrentio-scraper Torrentio is a Stremio addon ecosystem that scrapes public torrent providers and serves the results as Stremio stream results. The reposito… | 77 | 1280 | active |
| hadynz/obsidian-kindle-plugin An Obsidian plugin that syncs Kindle notes and highlights into your vault, either by screen-scraping Amazon's Kindle Reader cloud library o… | 75 | 1277 | active |
| sardanioss/httpcloak httpcloak is a Go HTTP client library that reproduces browser-identical TLS, HTTP/2, and HTTP/3 fingerprints (JA3/JA4, Akamai, header order… | 61 | 1276 | active |
| brunosimon/my-room-in-3d A 3D interactive recreation of the author's room built with Three.js and JavaScript, viewable in the browser. It is a creative demo/portfol… | 32 | 4486 | maintenance |
| acgotaku/YAAW-for-Chrome A Chrome extension providing a web frontend dashboard for the Aria2 download manager via its JSON-RPC interface. It can intercept browser d… | 71 | 1270 | active |
| kkangert/kspider Kspider is a self-hosted visual web scraping platform written in Java where users define crawler workflows as flowcharts without writing ba… | 14 | 1269 | active |
| umutxyp/MusicBot Beatra is an advanced Discord music bot built on discord.js v14 that streams music from YouTube, Spotify, SoundCloud, and direct links with… | 70 | 1266 | active |
| mvdan/xurls A Go library and CLI tool that extracts URLs from arbitrary text using regular expressions built from TLD lists. It offers Relaxed and Stri… | 67 | 1265 | active |
| cclank/news-aggregator-skill A Python-based agent skill that aggregates news from 44+ sources (tech, finance, AI, international) and generates AI-summarized daily brief… | 53 | 1263 | active |
| 0xHJK/music-dl A Python 3 command line tool that aggregates search across multiple Chinese music sites (NetEase, QQ Music, Kugou, Baidu, Xiami, Migu) and … | 39 | 4436 | maintenance |
| l429609201/misaka_danmu_server A self-hosted danmaku (bullet comment) aggregation and management server written in Python, compatible with the dandanplay API specificatio… | 81 | 1256 | active |
| firecrawl/fire-enrich Fire Enrich is an AI-powered data enrichment web application that transforms a list of email addresses into rich company datasets, includin… | 39 | 1255 | active |
| myreader-io/myGPTReader myGPTReader is a Slack bot powered by ChatGPT that reads and summarizes webpages, documents (eBooks, PDF, DOCX), and YouTube videos, and su… | 60 | 4419 | maintenance |
| JustinBeckwith/linkinator Linkinator is a broken link checker that crawls websites, local HTML files, and markdown documentation to find dead links and invalid URLs.… | 99 | 1254 | active |
| MarioVilas/googlesearch A Python library that performs Google web searches programmatically and returns result URLs, without using an official Google API. It is un… | 10 | 1252 | active |
| minsight-ai-info/AI-Search-Hub AI Search Hub is an open-source Skill that aggregates native AI search capabilities from platforms like Gemini, Grok, Doubao, and Yuanbao i… | 50 | 1250 | active |
| EmergenceAI/Agent-E Agent-E is an open-source agent-based system that automates actions on the user's computer, currently focusing on browser automation via na… | 61 | 1248 | active |
| egbertbouman/youtube-comment-downloader A Python script and library for downloading YouTube video comments without using the official YouTube API. It outputs comments in JSONL, JS… | 88 | 1247 | active |
| AIPexStudio/AIPex AIPex is an open-source (MIT) AI browser automation assistant delivered as a Chrome/Edge extension that runs inside your existing browser w… | 81 | 1244 | active |
| anvaka/npmgraph.an A web application that visualizes the dependency graph of npm packages in 2D and 3D, built with Vue 3 and the ngraph library. It fetches pa… | 66 | 1243 | active |
| tomorrow505/auto_feed_js A Tampermonkey/Greasemonkey userscript that enables one-click cross-posting (re-seeding) of torrents across 100+ private tracker sites. It … | 77 | 1242 | active |
| nonbili/NouTube NouTube is an Android and desktop application that wraps the mobile YouTube and YouTube Music web apps in a webview, adding ad blocking, ba… | 81 | 1239 | active |
| Decodo/Decodo Decodo (formerly Smartproxy) is a commercial rotating proxy network and web scraping platform offering 125M+ residential, mobile, ISP, and … | 69 | 1238 | active |
| jeremykendall/php-domain-parser A PHP library that parses domain names into their component parts (subdomain, registrable domain, second-level domain, public suffix) using… | 55 | 1238 | active |
| seveniruby/AppCrawler AppCrawler is a Scala-based automated app traversal (crawler) testing tool built on Appium, supporting Android, iOS, mini-programs, and Har… | 58 | 1236 | active |
| raawaa/jav-scrapy A TypeScript-based Node.js CLI tool that batch-scrapes JAV (adult video) metadata, magnet links, and cover images from source websites. It … | 94 | 1235 | active |
| eatmoreduck/boss-zhipin-scraper A Python CLI scraper for BOSS Zhipin (zhipin.com) that connects to a locally logged-in Chrome via the Chrome DevTools Protocol to call the … | 57 | 1231 | active |
| nicobailon/pi-web-access A TypeScript extension for the Pi coding agent that adds web search, URL content extraction, GitHub repo cloning, PDF extraction, and YouTu… | 83 | 1229 | active |
| danny0838/webscrapbook WebScrapBook is a browser extension that captures web pages faithfully into various archive formats (including MAFF and HTZ) for local or b… | 77 | 1228 | active |
| iszhouhua/social-media-copilot An open-source browser extension (built with WXT and TypeScript) that scrapes data from Chinese social media platforms including Xiaohongsh… | 65 | 1228 | active |
| saeeddhqan/Maryam OWASP Maryam is a modular open-source OSINT framework for harvesting data from open sources, search engines, and social networks. It provid… | 10 | 1228 | active |
| ImAiiR/QobuzDownloaderX QobuzDownloaderX (QBDLX) is a desktop GUI application that downloads music streams directly from the Qobuz streaming platform using its API… | 78 | 1226 | active |
| daijro/browserforge BrowserForge is a Python library that generates realistic browser headers and fingerprints, mimicking real-world browser, OS, and device di… | 56 | 1224 | active |
| eooce/Auto-login-netlib A JavaScript automation script that periodically logs into netlib.re accounts to keep free domains alive, run via GitHub Actions on a 60-da… | 39 | 1224 | active |
| firecrawl/web-agent An open-source TypeScript framework for building autonomous web research agents, layered from a Next.js chat template down to an agent core… | 50 | 1222 | active |
| mika-cn/maoxian-web-clipper MaoXian Web Clipper is a browser extension for Firefox and Chrome that clips selected content from web pages and saves it to the local mach… | 67 | 1221 | active |