function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Threezh1/JSFinder JSFinder is a Python command-line tool that crawls a website's JavaScript files and extracts URLs and subdomains using regex parsing. It su… | 32 | 2976 | maintenance |
| hongchacha/cartoon An Android app dedicated to reading comics and manga, aggregating content from dozens of Chinese comic websites. It replaces a web browser … | 32 | 2952 | maintenance |
| facundoolano/google-play-scraper A Node.js library that scrapes application data from the Google Play store, exposing methods for app details, search, reviews, permissions,… | 66 | 2949 | maintenance |
| EvilCult/iptv-m3u-maker A Python tool that scrapes publicly shared IPTV live-stream sources, tests each link's latency on the local network, and generates an optim… | 32 | 2890 | maintenance |
| adieuadieu/serverless-chrome Serverless Chrome is a scaffolding framework and Serverless Framework plugin for running headless Chrome/Chromium on AWS Lambda. It bundles… | 23 | 2887 | maintenance |
| howie6879/owllook owllook is a self-hosted vertical search engine for Chinese web novels, built on Python with Sanic, MongoDB, and Redis. It aggregates resul… | 32 | 2871 | maintenance |
| CharlesPikachu/DecryptLogin A Python library providing programmatic login APIs for popular websites (Weibo, Bilibili, Zhihu, GitHub, Taobao, etc.) built on the request… | 32 | 2855 | maintenance |
| hello-efficiency-inc/raven-reader Raven Reader is an open-source cross-platform desktop RSS/news reader built with Electron and Vue that aggregates articles from feeds in a … | 10 | 2810 | maintenance |
| ganlvtech/down_52pojie_cn A single-page file explorer built with Vue.js that can be hosted on any static website, originally created to power the 52pojie forum's dow… | 23 | 2806 | maintenance |
| topfunky/hpple Hpple is an Objective-C wrapper around the XPathQuery libxml2 library for parsing HTML and searching it with XPath expressions. It was insp… | 32 | 2799 | maintenance |
| bhavsec/reconspider ReconSpider is an open-source OSINT framework written in Python for scanning IP addresses, emails, websites, and organizations to gather in… | 23 | 2778 | maintenance |
| yann-shi/dht A Go library implementing the BitTorrent DHT protocol (BEP-3, 5, 9, 10) with two modes: a standard DHT server and a crawling mode for harve… | 23 | 2769 | maintenance |
| brianway/webporter webporter is a Java crawler application built on the webmagic framework that demonstrates a complete pipeline of data crawling, persistence… | 23 | 2766 | maintenance |
| DormyMo/SpiderKeeper SpiderKeeper is a self-hosted web-based admin dashboard for managing Scrapy spiders running on Scrapyd servers. It provides spider scheduli… | 32 | 2763 | maintenance |
| Xmader/musescore-downloader A tool for downloading sheet music from musescore.com for free without a login or Musescore Pro subscription, available as a CLI (npx msdl)… | 23 | 2762 | maintenance |
| jeanphix/Ghost.py Ghost.py is a scriptable WebKit-based web client library for Python, built on PySide2/Qt5, allowing programmatic page loading and content i… | 32 | 2756 | maintenance |
| martinvigo/email2phonenumber A Python OSINT tool that discovers a target's phone number from just their email address by abusing password reset flows that leak masked p… | 32 | 2747 | maintenance |
| jerry3747/taobao_seckill A Python Selenium-based script that automates flash-sale purchases on Taobao and Tmall, timing checkout to grab discounted items like TVs o… | 32 | 2689 | maintenance |
| pocketjoso/penthouse Penthouse is a Node.js library that generates critical path CSS for web pages, extracting the CSS needed to render above-the-fold content f… | 23 | 2680 | maintenance |
| mckaywrigley/paul-graham-gpt An open-source AI-powered search and chat application built over Paul Graham's essays using OpenAI embeddings, pgvector on Supabase, and GP… | 30 | 2663 | maintenance |
| weskerfoot/DeleteFB A Python CLI tool that uses Selenium to automate deleting your Facebook posts, page likes, and conversations through your real browser. It … | 10 | 2661 | maintenance |
| thewhiteh4t/nexfil Nexfil is an OSINT command-line tool written in Python that finds social media profiles by username across 350+ websites in seconds. It sup… | 32 | 2610 | maintenance |
| mattt/Ono Ono is an Objective-C/Swift library for parsing and querying XML and HTML documents on iOS and macOS, built on libxml2. It provides a conve… | 32 | 2595 | maintenance |
| Ccixyj/JBusDriver An Android app for browsing the JAVBUS adult content catalog, inspired by JAViewer. It is built with Kotlin using MVP architecture, RxJava2… | 23 | 2585 | maintenance |
| obheda12/GitDorker GitDorker is a Python CLI tool that uses the GitHub Search API with a curated list of over 200 dorks to find sensitive information exposed … | 32 | 2577 | maintenance |
| Tuhinshubhra/CMSeeK CMSeeK is a Python 3 command-line suite that detects the content management system (CMS) powering a website, supporting over 180 CMSs inclu… | 65 | 2571 | maintenance |
| justinzm/gopup GoPUP is a Python library that provides convenient interfaces to a wide range of public Chinese data sources, including Baidu/Weibo/Google … | 23 | 2548 | maintenance |
| luin/readability A Node.js library that extracts clean, readable article content from any web page, based on arc90's readability project. It returns the art… | 32 | 2519 | maintenance |
| xtuhcy/gecco Gecco is a lightweight, easy-to-use web crawler framework for Java that lets developers define crawlers with annotation-based jQuery-style … | 51 | 2510 | maintenance |
| yichahucha/surge A collection of JavaScript rewrite scripts for the Surge and Quantumult X proxy apps. They modify HTTP responses on iOS to inject Netflix I… | 70 | 2468 | maintenance |
| lorien/grab Grab is a Python web scraping framework providing HTTP request handling, proxy/cookie support, and XPath-based HTML parsing, plus a Spider … | 51 | 2463 | maintenance |
| andyzys/jd_seckill A Python script that automates logging into JD.com (Jingdong), reserving products like Moutai liquor, and purchasing them at a scheduled se… | 32 | 2450 | maintenance |
| Anning01/AIMedia AIMedia is a heavyweight integrated application that automatically scrapes trending news, generates AI-written articles with AI-generated i… | 63 | 2436 | maintenance |
| decaywood/XueQiuSuperSpider A Java 8 web scraping framework for collecting stock data from Xueqiu (Snowball) and other Chinese financial sites. It is built around comp… | 32 | 2433 | maintenance |
| netlify/staticgen StaticGen.com is a Gatsby-built leaderboard website that ranks open-source static site generators by GitHub/GitLab stars, forks, and Twitte… | 10 | 2433 | maintenance |
| screetsec/Sudomy Sudomy is a Bash-based subdomain enumeration and reconnaissance framework that collects subdomains via active brute-forcing and passive thi… | 23 | 2432 | maintenance |
| delight-im/Android-AdvancedWebView An enhanced WebView component library for Android that fixes common WebView quirks and works out of the box. It provides a drop-in replacem… | 32 | 2412 | maintenance |
| paquettg/php-html-parser A PHP library that parses HTML into a DOM and lets you find and manipulate tags using CSS selectors, similar to jQuery. It is designed for … | 32 | 2398 | maintenance |
| madelynnblue/goread Go Read is a Google Reader clone: a self-hostable RSS feed reader built in Go on Google App Engine with an AngularJS frontend. It lets user… | 10 | 2364 | maintenance |
| QianyanTech/Image-Downloader A Python application that crawls and downloads images from Google, Bing, and Baidu using Selenium or API drivers. It offers both a PyQt5 GU… | 23 | 2362 | maintenance |
| tangrams/tangram Tangram is a JavaScript library for rendering 2D and 3D maps live in the browser using WebGL, tuned for OpenStreetMap but supporting GeoJSO… | 52 | 2334 | maintenance |
| hidroh/materialistic Materialistic is a material-design Hacker News client for Android built in Java. It uses the official Hacker News API, Algolia search, and … | 23 | 2334 | maintenance |
| Encom Boardroom A WebGL/Three.js web application recreating the Encom boardroom scene from Tron: Legacy, visualizing realtime GitHub and Wikipedia data str… | 32 | 2318 | maintenance |
| lucasjinreal/weibo_terminater A Python-based web scraper that crawls Weibo (Sina's microblog platform) to collect user posts, comments, followers, and conversation pairs… | 32 | 2317 | maintenance |
| ggerganov/wave-share A proof-of-concept browser application that shares files between nearby devices by performing WebRTC signaling through audio tones instead … | 32 | 2311 | maintenance |
| Terminus2049/Terminus2049.github.io Terminus (端点星) is a GitHub Pages-hosted site built with Jekyll that archives articles deleted from Chinese platforms like WeChat and Weibo … | 32 | 2308 | maintenance |
| sfu-db/dataprep DataPrep is a Python library for low-code data preparation, offering modules to collect data from common APIs (connector), run fast explora… | 23 | 2248 | maintenance |
| buckket/twtxt twtxt is a decentralised, minimalist microblogging tool and format specification where your identity is a publicly accessible text file of … | 23 | 2236 | maintenance |
| j0be/PowerDeleteSuite A browser bookmarklet for Reddit that lets users mass-edit and mass-delete their own comments and submissions using Reddit's API endpoints.… | 23 | 2173 | maintenance |
| ageitgey/node-unfluff A Node.js library and CLI tool that automatically extracts the main body content and metadata (title, author, date, images, tags, links) fr… | 32 | 2158 | maintenance |
| t9tio/cloudquery CloudQuery is a tool that turns any website into a JSON API by fetching pages with headless Chrome and extracting data via CSS selectors. I… | 32 | 2149 | maintenance |
| php-embed/Embed A PHP library that extracts metadata and embed information from any web page or web service using oEmbed, OpenGraph, Twitter Cards, and HTM… | 93 | 2140 | maintenance |
| jadepeng/XMusicDownloader A C# desktop application that aggregates music search across Baidu, NetEase, QQ, Kugou, and Migu music sites and supports batch downloading… | 23 | 2140 | maintenance |
| anouarbensaad/vulnx VulnX is a Python CLI tool that detects CMS types (WordPress, Joomla, Drupal, etc.), gathers target information like subdomains and DNS rec… | 23 | 2138 | maintenance |
| D4Vinci/Cr3dOv3r Cr3dOv3r is a Python command-line pentesting tool for investigating credential reuse attacks. Given an email, it searches public breach dat… | 57 | 2137 | maintenance |
| initstring/cloud_enum A Python command-line OSINT tool that enumerates publicly exposed resources across AWS, Azure, and Google Cloud using keyword mutations and… | 79 | 2132 | maintenance |
| UnaPibaGeek/ctfr CTFR is a Python command-line tool that enumerates HTTPS website subdomains by querying Certificate Transparency logs (via crt.sh) instead … | 32 | 2117 | maintenance |
| cyberagiinc/DevDocs DevDocs is a free, private, UI-based MCP server that crawls and extracts technical documentation (using Crawl4AI and Playwright) and expose… | 50 | 2106 | maintenance |
| TideSec/WDScanner WDScanner is a self-hosted distributed web vulnerability scanning platform written in PHP with Python backend workers. It combines customer… | 32 | 2100 | maintenance |
| iSafeBlue/TrackRay TrackRay (溯光) is an open-source penetration testing framework written in Java on SpringBoot that implements its own vulnerability scanning … | 23 | 2078 | maintenance |
| stevenvachon/broken-link-checker A Node.js library and CLI tool (blc) that crawls HTML to find broken links, missing images, and other dead URLs. It supports concurrent, st… | 32 | 2075 | maintenance |
| vim-awesome/vim-awesome Vim Awesome is an open-source web application that serves as a comprehensive directory of Vim plugins, aggregating data from GitHub, Vim.or… | 23 | 2063 | maintenance |
| Aabyss-Team/ARL ARL (Asset Reconnaissance Lighthouse) is a self-hosted asset reconnaissance system that quickly discovers internet-facing assets associated… | 29 | 2055 | maintenance |
| PuerkitoBio/gocrawl gocrawl is a polite, slim and concurrent web crawler library written in Go. It respects robots.txt rules, applies per-host crawl delays, an… | 23 | 2052 | maintenance |
| 670848654/SakuraAnime A third-party Android client for the anime streaming sites Yhdm (Sakura Anime) and SiliSili, built in Java using jsoup for scraping site co… | 10 | 2043 | maintenance |
| 0xbug/Hawkeye Hawkeye is a self-hosted system that monitors GitHub for leaked sensitive information, such as employees pushing company code or credential… | 23 | 2033 | maintenance |
| Nekmo/dirhunt Dirhunt is a Python CLI web crawler optimized for finding and analyzing web directories without brute-forcing paths. It detects 'index of' … | 23 | 2007 | maintenance |
| bigemon/ChatGPT-ToolBox A ChatGPT toolbox userscript/browser extension written largely by ChatGPT itself, injectable via Tampermonkey or Chrome bookmarklets. It ad… | 32 | 2004 | maintenance |
| awolfly9/IPProxyTool A Python/Scrapy application that crawls free proxy websites, validates the collected proxy IPs against target sites, and stores usable prox… | 32 | 1998 | maintenance |
| youusername/magnetX magnetX is a native macOS application written in Objective-C for searching magnet links across multiple torrent search sites without browse… | 23 | 1984 | maintenance |
| microsoft/magentic-ui MagenticLite is an experimental agentic application from Microsoft AI Frontiers that automates tasks across the browser and local file syst… | 81 | 10077 | experimental |
| vaguileradiaz/tinfoleak tinfoleak is an open-source Python tool for OSINT/SOCMINT analysis of Twitter accounts, extracting structured intelligence such as user act… | 32 | 1980 | maintenance |
| coursera-dl/edx-dl A Python command-line tool that downloads video lectures and course materials from Open edX-based platforms such as edX.org, Stanford Onlin… | 23 | 1962 | maintenance |
| D35m0nd142/LFISuite LFISuite is a fully automatic Python tool that scans for and exploits Local File Inclusion (LFI) vulnerabilities using eight different atta… | 23 | 1961 | maintenance |
| SathyaBhat/spotify-dl A Python CLI tool that fetches track metadata from Spotify playlists, albums, or tracks via the Spotify Web API and downloads the correspon… | 64 | 1944 | maintenance |
| Xyntax/POC-T POC-T is a Python 2.7 plugin-based concurrent framework for penetration testing tasks such as crawling, bruteforcing, and batch PoC/EXP ver… | 23 | 1936 | maintenance |
| stevenyomi/copymanga An unofficial Tachiyomi extension that adds CopyManga (拷贝漫画) as a manga reading source. It is installed as a trusted extension within the T… | 10 | 1919 | maintenance |
| bugswriter/notflix Notflix is a shell script that scrapes 1337x for magnet links and streams the resulting torrents via peerflix. It is a lightweight command-… | 32 | 1906 | maintenance |
| scrapy/scrapely Scrapely is a pure-Python library for extracting structured data from HTML pages. It learns a parser from example pages annotated with the … | 32 | 1883 | maintenance |
| eldraco/domain_analyzer Domain Analyzer is a Python-based security analysis tool that automatically discovers and reports information about a given domain, includi… | 32 | 1861 | maintenance |
| kanishka-linux/reminiscence Reminiscence is a self-hosted bookmark and archive manager built with Django that saves web pages in HTML, PDF, or full-page PNG formats. I… | 23 | 1855 | maintenance |
| alvarobartt/investpy investpy is a Python package for retrieving recent and historical financial data from Investing.com, covering stocks, funds, ETFs, indices,… | 57 | 1851 | maintenance |
| awake1t/linglong Linglong is a self-hosted asset reconnaissance and scanning system written in Go that continuously discovers network assets using masscan+n… | 32 | 1846 | maintenance |
| thedaviddelta/lingva-translate Lingva Translate is a free and open-source alternative front-end for Google Translate that supports over a hundred languages. It scrapes Go… | 32 | 1840 | maintenance |
| xianhu/PSpider PSpider is a simple, easy-to-read web spider framework written in Python 3.8+. It uses a Fetcher/Parser/Saver pipeline with queues for mult… | 32 | 1835 | maintenance |
| kickscondor/fraidycat Fraidycat is an app (and browser extension) for following people across many platforms - blogs, wikis, YouTube, Twitter, Reddit, Instagram … | 23 | 1818 | maintenance |
| chaychan/TouTiao An Android app that closely replicates the Toutiao (Today's Headlines) news client, built with RxJava, Retrofit, and MVP architecture. It s… | 32 | 1813 | maintenance |
| Yvesssn/DetectDee DetectDee is a Go CLI tool for OSINT that hunts down social media accounts by username, email, or phone number across many social networks.… | 20 | 1806 | maintenance |
| vertex-app/vertex Vertex is a self-hosted management tool for private tracker (PT) users that combines TV-show tracking with automated torrent snatching ('sh… | 70 | 1792 | maintenance |
| anvaka/pm Software Galaxies is an interactive web visualization that renders major software package manager ecosystems (npm, Bower, Composer, RubyGem… | 72 | 1780 | maintenance |
| 1tayH/noisy A Python script that generates random HTTP and DNS traffic noise in the background to obscure your real browsing patterns. It crawls config… | 32 | 1778 | maintenance |
| megadose/OnionSearch OnionSearch is a Python 3 CLI script that scrapes URLs from multiple .onion search engines such as Ahmia, Phobos, and Deeplink. It supports… | 32 | 1777 | maintenance |
| reworkd/tarsier Tarsier is a Python library providing vision utilities for LLM-driven web interaction agents. It visually tags interactable page elements w… | 17 | 1761 | maintenance |
| htmlpreview/htmlpreview.github.com A client-side web tool that renders HTML files hosted on GitHub or BitBucket repositories without cloning or downloading them. It prepends … | 32 | 1759 | maintenance |
| intoli/remote-browser Remote Browser is a JavaScript library for programmatically controlling browsers like Chrome and Firefox, built on the standard Web Extensi… | 32 | 1750 | maintenance |
| DanMcInerney/xsscrapy A Python-based spider built on Scrapy that crawls a website and tests every link it finds for cross-site scripting (XSS) and basic SQL inje… | 32 | 1747 | maintenance |
| Dimillian/SwiftHN SwiftHN is an open-source Hacker News reader iOS app written in Swift, published on the App Store as HN Reader. It uses its own scraping li… | 32 | 1741 | maintenance |
| howie6879/ruia Ruia is an async web scraping micro-framework for Python 3.6+ built on asyncio and aiohttp. It offers declarative Item/Field extraction (XP… | 23 | 1738 | maintenance |
| iqiqiya/iqiqiya-API A PHP-based collection of free web API endpoints for parsing media from Chinese platforms (Douyin, Kuaishou, Bilibili, NetEase Music, Ximal… | 10 | 1737 | maintenance |
| BKcore/HexGL HexGL is a futuristic, fast-paced HTML5 racing game built with JavaScript and WebGL using three.js, inspired by Wipeout and F-Zero. This re… | 32 | 1736 | maintenance |