function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| MakiNaruto/Automatic_ticket_purchase A Python script that automates ticket purchasing on Damai (大麦网), China's major ticketing platform, using Selenium for login and requests-ba… | 32 | 5613 | abandoned |
| liuwons/wxBot A Python library for building WeChat bots using the web WeChat interface, supporting message handling and automated replies. The project is… | 32 | 5314 | abandoned |
| PeterDing/iScript A collection of Python 2 command-line scripts for downloading and playing media from Chinese services like Xiami, NetEase Music, Baidu Musi… | 32 | 5119 | abandoned |
| r0oth3x49/udemy-dl A cross-platform Python command-line utility for downloading Udemy course videos and subtitles for personal offline use. It supports resumi… | 10 | 4954 | abandoned |
| ohld/igbot A Python library and collection of scripts providing an unofficial Instagram API wrapper with bot features like auto-follow, auto-like, and… | 52 | 4876 | abandoned |
| 0x5e/wechat-deleted-friends A Python script that detects which WeChat contacts have deleted you by attempting to add them to a new group chat via the WeChat Web API. T… | 10 | 4774 | abandoned |
| hanc00l/wooyun_public A crawler and search application for the archived Wooyun.org security vulnerability disclosure platform, containing ~40k-88k public vulnera… | 10 | 4399 | abandoned |
| ruicky/jd_sign_bot A JD.com (Jingdong) sign-in bot that automates daily check-ins on the JD e-commerce platform. The project has been discontinued, with the R… | 10 | 4368 | abandoned |
| qiyeboy/IPProxyPool IPProxyPool is a Python proxy pool service that crawls free proxy IPs from the web, validates them, stores them in a database (SQLite by de… | 23 | 4286 | abandoned |
| Nemo2011/bilibili-api A Python SDK for programmatically accessing Bilibili (bilibili.com), including video data, danmaku comments, user info, and login. The repo… | 10 | 4188 | abandoned |
| Greenwolf/social_mapper Social Mapper is a Python 3 OSINT tool that enumerates and correlates social media profiles across sites like LinkedIn, Facebook, Twitter, … | 32 | 4073 | abandoned |
| bisguzar/twitter-scraper A Python library that scrapes Twitter's frontend JavaScript API without authentication, letting users fetch tweets from profiles or hashtag… | 10 | 4005 | abandoned |
| pyppeteer/pyppeteer Pyppeteer is an unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. It supports page navigati… | 32 | 3942 | abandoned |
| PokemonGoF/PokemonGo-Bot A community-developed bot for the game Pokemon Go that automates gameplay actions like spinning pokestops and catching pokemon. It is writt… | 23 | 3915 | abandoned |
| CouchPotato/CouchPotatoServer CouchPotato is a self-hosted Python application that automatically searches for and downloads movies via NZBs and torrents. Users maintain … | 10 | 3871 | abandoned |
| sqzw-x/mdcx MDCx is a Python-based desktop application and Docker-deployable service that scrapes movie metadata from multiple online sources, aggregat… | 94 | 3725 | abandoned |
| miyakogi/pyppeteer An unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. This original repository has moved to … | 10 | 3550 | abandoned |
| amir20/phantomjs-node A NodeJS wrapper providing a promise-based API around the PhantomJS headless browser. It lets JavaScript code create PhantomJS instances, o… | 10 | 3517 | abandoned |
| codeestX/GeekNews GeekNews is an Android reading app for programmers that aggregates content from Zhihu Daily, WeChat news, Gank.io, Juejin, and V2EX in a Ma… | 32 | 3491 | abandoned |
| bowenpay/wechat-spider A Python-based web crawler for scraping articles from WeChat public accounts (微信公众号), built on Django with MySQL and Redis, including a web… | 32 | 3369 | abandoned |
| LiuXingMing/SinaSpider A Python web crawler for Sina Weibo (Chinese microblog) built on Scrapy, with three versions: a standalone spider, a distributed version us… | 32 | 3285 | abandoned |
| gnemoug/distribute_crawler A distributed web crawler built on Scrapy, Redis, MongoDB, and Graphite, demonstrated with a spider for a Chinese book-download site. Redis… | 32 | 3238 | abandoned |
| harismuneer/Ultimate-Social-Scrapers A collection of Python-based scraping tools that extract public data from Facebook, Instagram, and Twitter (X), including posts, media, fol… | 44 | 3151 | abandoned |
| boramalper/magnetico magnetico is a self-hosted BitTorrent DHT search engine suite consisting of magneticod, an autonomous DHT crawler and metadata fetcher, and… | 10 | 3133 | abandoned |
| automagica/automagica Automagica was an open-source, AI-powered Robotic Process Automation (RPA) suite in Python, including a bot runtime, visual flow designer, … | 32 | 3099 | abandoned |
| airingursb/bilibili-user A Python web crawler that scrapes Bilibili user profiles (id, nickname, gender, avatar, level, birthday, location, etc.) and stores them in… | 32 | 3090 | abandoned |
| laurentj/slimerjs SlimerJS is a scriptable headless browser that provides the PhantomJS API on top of Gecko (Firefox) instead of WebKit, letting external Jav… | 23 | 2998 | abandoned |
| the0demiurge/ShadowSocksShare A Python/Flask web service that crawls shared Shadowsocks/ShadowsocksR accounts from public sharing websites, validates their connectivity,… | 10 | 2992 | abandoned |
| JAVClub/core A self-hosted adult video (JAV) library platform that automatically scrapes metadata, downloads videos via torrents, uploads them to Google… | 10 | 2875 | abandoned |
| NikolaiT/GoogleScraper GoogleScraper is a Python module and CLI tool for scraping search engine results from Google, Bing, Yandex, DuckDuckGo and others, with sup… | 32 | 2874 | abandoned |
| YahooArchive/anthelion Anthelion is an Apache Nutch plugin for focused crawling of semantic data embedded in HTML pages. It uses an online learning classifier to … | 10 | 2827 | abandoned |
| lanbing510/DouBanSpider A Python web scraper for Douban Books that crawls book listings by tag, storing ratings and review counts into Excel files. The author also… | 32 | 2786 | abandoned |
| spyglass-search/spyglass Spyglass is a cross-platform desktop personal search engine that crawls and indexes your local documents, saved web content, and connected … | 10 | 2721 | abandoned |
| Netflix-Skunkworks/Scumblr Scumblr is a Ruby on Rails web application from Netflix for performing periodic syncs of data sources (GitHub repos, URLs, DNS) and running… | 23 | 2642 | abandoned |
| loadchange/amemv-crawler A Python 3 script that downloads all videos from a specified Douyin (TikTok China) user account, as well as all videos under a given challe… | 32 | 2641 | abandoned |
| zqjzqj/mtSecKill A command-line tool for automatically grabbing (seckill) Moutai liquor purchases on JD.com (Jingdong). It automates the flash-sale checkout… | 32 | 2580 | abandoned |
| IvanGlinkin/CCTV CCTV (Close-Circuit Telegram Vision) is an open-source OSINT tool that abuses Telegram's 'People Nearby' feature to triangulate and track u… | 28 | 2478 | abandoned |
| taspinar/twitterscraper A Python library that scrapes tweets and user information from Twitter using requests and BeautifulSoup, without relying on Twitter's offic… | 23 | 2461 | abandoned |
| QiuChenlyOpenSource/MusicDownload A tool for downloading songs in FLAC/MP3 quality, primarily using the QQ Music API. The project has been discontinued after legal pressure … | 10 | 2449 | abandoned |
| scrapoxy/scrapoxy Scrapoxy was an open-source proxy manager for web scraping that aggregated proxies from cloud providers and other sources behind a single A… | 62 | 2414 | abandoned |
| Hari-Nagarajan/fairgame FairGame is a Python application that monitors Amazon for out-of-stock products and can automatically place orders when items become availa… | 23 | 2410 | abandoned |
| jaeles-project/jaeles Jaeles is a Go-based framework for building and running automated web application vulnerability scanners using customizable YAML signatures… | 62 | 2370 | abandoned |
| egrcc/zhihu-python A Python 2.7 library for scraping content from Zhihu, a Chinese Q&A platform, including questions, answers, users, and favorites. It can ex… | 32 | 2335 | abandoned |
| chiphuyen/lazynlp A Python library for crawling, cleaning, and deduplicating web pages to build massive monolingual text datasets, suitable for training lang… | 23 | 2284 | abandoned |
| PaulMcInnis/JobFunnel JobFunnel is a Python CLI tool that scrapes job postings from multiple job websites (Indeed, Glassdoor, LinkedIn) into a single deduplicate… | 10 | 2180 | abandoned |
| ckreibich/scholar.py A Python module that queries and parses Google Scholar search results, extracting publication metadata, citation counts, PDF links, and Bib… | 10 | 2175 | abandoned |
| UnkL4b/GitMiner GitMiner is a Python CLI tool for advanced searching and mining of code and code snippets on GitHub, often used to find sensitive informati… | 45 | 2152 | abandoned |
| minimaxir/facebook-page-post-scraper A Python script collection that scrapes all posts, reactions, and comments from public Facebook Pages and open Groups via the Facebook Grap… | 10 | 2135 | abandoned |
| simplecrawler/simplecrawler simplecrawler is a flexible, event-driven web crawler library for Node.js with a configurable queue system, robots.txt support, and link di… | 10 | 2134 | abandoned |
| brenden/node-webshot A Node.js library providing a simple API for taking website screenshots by wrapping PhantomJS's WebKit rendering. It supports capturing URL… | 32 | 2109 | abandoned |
| althonos/InstaLooter InstaLooter is a Python CLI tool that downloads pictures and videos from Instagram profiles without using the official API. It is a re-impl… | 23 | 2098 | abandoned |
| fouber/page-monitor A Node.js library that uses PhantomJS to render webpages, capture screenshots, and diff DOM changes (added/removed elements, text, and styl… | 32 | 2092 | abandoned |
| mukulhase/WebWhatsapp-Wrapper A Python library providing an unofficial API for WhatsApp by automating WhatsApp Web through Selenium browser automation. It allows sending… | 23 | 2077 | abandoned |
| yahoo/gryffin Gryffin is a large-scale web security scanning platform written in Go, built on a publisher-subscriber architecture for horizontal scaling.… | 10 | 2052 | abandoned |
| zythum/mama2 MAMA2 is a browser bookmarklet/plugin project that replaces Flash video players on Chinese video sites (Bilibili, Youku, Tudou, Sohu, etc.)… | 32 | 2040 | abandoned |
| anime-dl/anime-downloader A Python command-line tool for downloading and streaming anime episodes from various streaming sites and Nyaa. It supports batch episode do… | 23 | 2002 | abandoned |
| trevorlinton/webkit.js A pure JavaScript port of WebKit's WebCore compiled via Emscripten, capable of rendering HTML5, CSS3, and SVG to WebGL/Canvas contexts in b… | 32 | 1971 | abandoned |
| wkeeling/selenium-wire Selenium Wire extends Selenium's Python bindings to expose the HTTP/HTTPS requests and responses made by the browser, with APIs to inspect … | 10 | 1965 | abandoned |
| mozilla/fathom Fathom is a supervised-learning framework for recognizing and classifying parts of web pages, tagging DOM nodes with types and probabilitie… | 10 | 1963 | abandoned |
| gxvv/ex-baiduyunpan A Tampermonkey userscript that removes Baidu Netdisk's large-file download restrictions and enables batch copying of download links. It wor… | 10 | 1914 | abandoned |
| cycz/jdBuyMask A Python automation tool that monitored JD.com (Jingdong) for face mask stock during the COVID-19 pandemic and automatically placed purchas… | 32 | 1849 | abandoned |
| hu17889/go_spider go_spider is a concurrent web crawler framework written in Go, designed for crawling vertical communities with a flexible, modular architec… | 23 | 1818 | abandoned |
| NotJoeMartinez/yt-fts yt-fts is a Python command line tool that downloads all subtitles from a YouTube channel or playlist via yt-dlp and stores them in a SQLite… | 59 | 1812 | abandoned |
| eth0izzle/bucket-stream A Python CLI tool that monitors certificate transparency logs via certstream and discovers public Amazon S3 buckets by generating permutati… | 36 | 1808 | abandoned |
| timschneeb/tachiyomi-extensions-archive A historical archive of removed manga source extensions for the Tachiyomi Android reader app. The repository was DMCA'd and erased, and is … | 10 | 1797 | abandoned |
| node-js-libs/node.io node.io is a Node.js web scraping and data extraction library originally written in 2010. It is explicitly no longer maintained, with the a… | 32 | 1792 | abandoned |
| bughandler/cnki-downloader A small desktop tool for searching and downloading academic literature from CNKI (China National Knowledge Infrastructure). Its backend int… | 48 | 1762 | abandoned |
| alex/nyt-2020-election-scraper A git-scraping application that periodically scrapes the New York Times' 2020 election results JSON API and commits snapshots to the reposi… | 32 | 1756 | abandoned |
| mchristopher/PokemonGo-DesktopMap An Electron desktop application that bundles the PokemonGo-Map project with an HTML UI and Python dependencies to show a live visualization… | 10 | 1747 | abandoned |
| metafates/mangal Mangal is a cross-platform CLI/TUI manga downloader written in Go, with built-in sources (Mangadex, Manganelo, Manganato, Mangapill), exten… | 10 | 1745 | abandoned |
| sorenlouv/fb-sleep-stats A Node.js proof-of-concept tool that polls Facebook's online/offline status to infer and visualize friends' sleep patterns. It consists of … | 32 | 1707 | abandoned |
| MShawon/YouTube-Viewer A Python-based multithreaded YouTube view bot that uses Selenium-driven browser sessions with free, premium, and rotating proxies to inflat… | 23 | 1663 | abandoned |
| cool2528/baiduCDP BaiduCDP is a C++ Windows application for high-speed downloading from Baidu Netdisk (Baidu Cloud Drive). It analyzes Baidu Netdisk's web AP… | 10 | 1661 | abandoned |
| tmort/Socialite Socialite is a small vanilla JavaScript library for lazily loading social sharing widgets (Facebook, Twitter, Google+, LinkedIn, Pinterest,… | 10 | 1660 | abandoned |
| sc1341/InstagramOSINT A Python CLI tool and importable module that scrapes publicly available profile information from Instagram accounts, such as follower count… | 10 | 1651 | abandoned |
| ZFC-Digital/puppeteer-real-browser A Node.js library that wraps Puppeteer with a real-browser profile to bypass bot detection systems like Cloudflare and Turnstile captchas. … | 35 | 1642 | abandoned |
| junhoyeo/threads-api An unofficial, reverse-engineered Node.js/TypeScript client for Meta's Threads social network, built shortly after Threads launched, with a… | 10 | 1621 | abandoned |
| erma0/douyin A Python crawler for Douyin (Chinese TikTok) that collected public data such as account profiles, likes, favorites, music, hashtags, search… | 76 | 1602 | abandoned |
| chinoogawa/fbht A Python 2 command-line tool for interacting with and scraping Facebook accounts, including graph-based analysis of social connections. It … | 23 | 1591 | abandoned |
| fossasia/loklak_wok_android Loklak Wok is an Android app that acts as a harvesting peer for the loklak_server, collecting social media messages (tweets) and pushing th… | 10 | 1558 | abandoned |
| fossasia/searss A small Python CLI tool that scrapes search results from engines like Google, Bing, DuckDuckGo, and Ask.com and converts them into RSS feed… | 10 | 1534 | abandoned |
| GravityLabs/goose Goose is a Scala library (originally Java) that extracts the main body text, metadata, publish date, embedded videos, and top image from ne… | 10 | 1526 | abandoned |
| fossasia/loklak-webtweets A small client-side web app that fetches and displays tweets via the loklak API, embeddable on any website with configurable query, count, … | 10 | 1525 | abandoned |
| fossasia/loklak_tweetheatmap A web application that displays a geographic heat map of tweets matching a search query, built with Angular.js and OpenLayers 3 using the L… | 32 | 1518 | abandoned |
| benitoro/stockholm A Python framework that crawls Shanghai and Shenzhen A-share stock market data from Yahoo YQL and Sina Finance, and tests stock-picking str… | 32 | 1517 | abandoned |
| githubwing/GankClient-Kotlin An Android client app for gank.io (a Chinese tech content aggregator) written in Kotlin with Material Design. It serves as a reference impl… | 32 | 1515 | abandoned |
| yhat/scrape A Go library providing a higher-level interface over golang.org/x/net/html for web scraping. It offers generic tree traversal helpers like … | 10 | 1515 | abandoned |
| FOSSASIA-Web/timeline.api.fossasia.net A jQuery plugin that embeds a FOSSASIA community event timeline into any website, fetching events from the FOSSASIA/Freifunk Calendar API o… | 32 | 1514 | abandoned |
| isanchop/stuhack A Chrome extension that attempted to unlock Studocu premium features such as removing premium banners, bypassing document blur, and enablin… | 10 | 1513 | abandoned |
| fossasia/wp-tweet-legacy A WordPress widget plugin that displays Twitter feeds in widget-ready areas, parsing @usernames, #hashtags, and URLs into links. It support… | 32 | 1504 | abandoned |
| qinxuye/cola Cola is a high-level distributed crawling framework in Python for scraping pages and extracting structured data from websites. The same cra… | 10 | 1499 | abandoned |
| fossasia/loklak_heatmapper A plain HTML/JavaScript web application that visualizes recent tweet activity as a heatmap on a world map. It fetches tweets via the loklak… | 10 | 1497 | abandoned |
| artyshko/smd A Python application that downloads high-quality music from Spotify playlists, albums, and songs by fetching audio from sources like YouTub… | 23 | 1486 | abandoned |
| johntitus/node-horseman A Node.js library providing a chainable, Promise-based API for controlling the PhantomJS headless browser, supporting page navigation, form… | 32 | 1485 | abandoned |
| yucccc/vue-mall A full-stack e-commerce demo mall (clone of the Smartisan/锤子 online store) built with Vue 2 + Vuex on the frontend and Node.js + MongoDB fo… | 32 | 1482 | abandoned |
| ring04h/wydomain wydomain is a Python command-line tool for discovering subdomains of a target domain. It combines dictionary-based DNS bruteforcing with qu… | 32 | 1479 | abandoned |
| fossasia/loklak_scraper_js A collection of JavaScript scrapers for the loklak project that extract data from websites like Twitter and Quora and output JSON resemblin… | 10 | 1473 | abandoned |
| fossasia/loklak-timeline-plugin A lightweight JavaScript plugin for embedding loklak search timelines into web pages via a simple HTML tag with data attributes. It renders… | 10 | 1457 | abandoned |
| keenwon/antcolony AntColony is a Node.js-based BitTorrent DHT network crawler that collects active infohashes, downloads and parses torrent files, and stores… | 32 | 1456 | abandoned |
| Entromorgan/Autoticket A Python/Selenium automation tool that automatically snipes and purchases tickets on Damai.cn (China's major ticketing site) by driving Chr… | 10 | 1450 | abandoned |