Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
MakiNaruto/Automatic_ticket_purchase
A Python script that automates ticket purchasing on Damai (大麦网), China's major ticketing platform, using Selenium for login and requests-ba…
325613abandoned
liuwons/wxBot
A Python library for building WeChat bots using the web WeChat interface, supporting message handling and automated replies. The project is…
325314abandoned
PeterDing/iScript
A collection of Python 2 command-line scripts for downloading and playing media from Chinese services like Xiami, NetEase Music, Baidu Musi…
325119abandoned
r0oth3x49/udemy-dl
A cross-platform Python command-line utility for downloading Udemy course videos and subtitles for personal offline use. It supports resumi…
104954abandoned
ohld/igbot
A Python library and collection of scripts providing an unofficial Instagram API wrapper with bot features like auto-follow, auto-like, and…
524876abandoned
0x5e/wechat-deleted-friends
A Python script that detects which WeChat contacts have deleted you by attempting to add them to a new group chat via the WeChat Web API. T…
104774abandoned
hanc00l/wooyun_public
A crawler and search application for the archived Wooyun.org security vulnerability disclosure platform, containing ~40k-88k public vulnera…
104399abandoned
ruicky/jd_sign_bot
A JD.com (Jingdong) sign-in bot that automates daily check-ins on the JD e-commerce platform. The project has been discontinued, with the R…
104368abandoned
qiyeboy/IPProxyPool
IPProxyPool is a Python proxy pool service that crawls free proxy IPs from the web, validates them, stores them in a database (SQLite by de…
234286abandoned
Nemo2011/bilibili-api
A Python SDK for programmatically accessing Bilibili (bilibili.com), including video data, danmaku comments, user info, and login. The repo…
104188abandoned
Greenwolf/social_mapper
Social Mapper is a Python 3 OSINT tool that enumerates and correlates social media profiles across sites like LinkedIn, Facebook, Twitter, …
324073abandoned
bisguzar/twitter-scraper
A Python library that scrapes Twitter's frontend JavaScript API without authentication, letting users fetch tweets from profiles or hashtag…
104005abandoned
pyppeteer/pyppeteer
Pyppeteer is an unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. It supports page navigati…
323942abandoned
PokemonGoF/PokemonGo-Bot
A community-developed bot for the game Pokemon Go that automates gameplay actions like spinning pokestops and catching pokemon. It is writt…
233915abandoned
CouchPotato/CouchPotatoServer
CouchPotato is a self-hosted Python application that automatically searches for and downloads movies via NZBs and torrents. Users maintain …
103871abandoned
sqzw-x/mdcx
MDCx is a Python-based desktop application and Docker-deployable service that scrapes movie metadata from multiple online sources, aggregat…
943725abandoned
miyakogi/pyppeteer
An unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. This original repository has moved to …
103550abandoned
amir20/phantomjs-node
A NodeJS wrapper providing a promise-based API around the PhantomJS headless browser. It lets JavaScript code create PhantomJS instances, o…
103517abandoned
codeestX/GeekNews
GeekNews is an Android reading app for programmers that aggregates content from Zhihu Daily, WeChat news, Gank.io, Juejin, and V2EX in a Ma…
323491abandoned
bowenpay/wechat-spider
A Python-based web crawler for scraping articles from WeChat public accounts (微信公众号), built on Django with MySQL and Redis, including a web…
323369abandoned
LiuXingMing/SinaSpider
A Python web crawler for Sina Weibo (Chinese microblog) built on Scrapy, with three versions: a standalone spider, a distributed version us…
323285abandoned
gnemoug/distribute_crawler
A distributed web crawler built on Scrapy, Redis, MongoDB, and Graphite, demonstrated with a spider for a Chinese book-download site. Redis…
323238abandoned
harismuneer/Ultimate-Social-Scrapers
A collection of Python-based scraping tools that extract public data from Facebook, Instagram, and Twitter (X), including posts, media, fol…
443151abandoned
boramalper/magnetico
magnetico is a self-hosted BitTorrent DHT search engine suite consisting of magneticod, an autonomous DHT crawler and metadata fetcher, and…
103133abandoned
automagica/automagica
Automagica was an open-source, AI-powered Robotic Process Automation (RPA) suite in Python, including a bot runtime, visual flow designer, …
323099abandoned
airingursb/bilibili-user
A Python web crawler that scrapes Bilibili user profiles (id, nickname, gender, avatar, level, birthday, location, etc.) and stores them in…
323090abandoned
laurentj/slimerjs
SlimerJS is a scriptable headless browser that provides the PhantomJS API on top of Gecko (Firefox) instead of WebKit, letting external Jav…
232998abandoned
the0demiurge/ShadowSocksShare
A Python/Flask web service that crawls shared Shadowsocks/ShadowsocksR accounts from public sharing websites, validates their connectivity,…
102992abandoned
JAVClub/core
A self-hosted adult video (JAV) library platform that automatically scrapes metadata, downloads videos via torrents, uploads them to Google…
102875abandoned
NikolaiT/GoogleScraper
GoogleScraper is a Python module and CLI tool for scraping search engine results from Google, Bing, Yandex, DuckDuckGo and others, with sup…
322874abandoned
YahooArchive/anthelion
Anthelion is an Apache Nutch plugin for focused crawling of semantic data embedded in HTML pages. It uses an online learning classifier to …
102827abandoned
lanbing510/DouBanSpider
A Python web scraper for Douban Books that crawls book listings by tag, storing ratings and review counts into Excel files. The author also…
322786abandoned
spyglass-search/spyglass
Spyglass is a cross-platform desktop personal search engine that crawls and indexes your local documents, saved web content, and connected …
102721abandoned
Netflix-Skunkworks/Scumblr
Scumblr is a Ruby on Rails web application from Netflix for performing periodic syncs of data sources (GitHub repos, URLs, DNS) and running…
232642abandoned
loadchange/amemv-crawler
A Python 3 script that downloads all videos from a specified Douyin (TikTok China) user account, as well as all videos under a given challe…
322641abandoned
zqjzqj/mtSecKill
A command-line tool for automatically grabbing (seckill) Moutai liquor purchases on JD.com (Jingdong). It automates the flash-sale checkout…
322580abandoned
IvanGlinkin/CCTV
CCTV (Close-Circuit Telegram Vision) is an open-source OSINT tool that abuses Telegram's 'People Nearby' feature to triangulate and track u…
282478abandoned
taspinar/twitterscraper
A Python library that scrapes tweets and user information from Twitter using requests and BeautifulSoup, without relying on Twitter's offic…
232461abandoned
QiuChenlyOpenSource/MusicDownload
A tool for downloading songs in FLAC/MP3 quality, primarily using the QQ Music API. The project has been discontinued after legal pressure …
102449abandoned
scrapoxy/scrapoxy
Scrapoxy was an open-source proxy manager for web scraping that aggregated proxies from cloud providers and other sources behind a single A…
622414abandoned
Hari-Nagarajan/fairgame
FairGame is a Python application that monitors Amazon for out-of-stock products and can automatically place orders when items become availa…
232410abandoned
jaeles-project/jaeles
Jaeles is a Go-based framework for building and running automated web application vulnerability scanners using customizable YAML signatures…
622370abandoned
egrcc/zhihu-python
A Python 2.7 library for scraping content from Zhihu, a Chinese Q&A platform, including questions, answers, users, and favorites. It can ex…
322335abandoned
chiphuyen/lazynlp
A Python library for crawling, cleaning, and deduplicating web pages to build massive monolingual text datasets, suitable for training lang…
232284abandoned
PaulMcInnis/JobFunnel
JobFunnel is a Python CLI tool that scrapes job postings from multiple job websites (Indeed, Glassdoor, LinkedIn) into a single deduplicate…
102180abandoned
ckreibich/scholar.py
A Python module that queries and parses Google Scholar search results, extracting publication metadata, citation counts, PDF links, and Bib…
102175abandoned
UnkL4b/GitMiner
GitMiner is a Python CLI tool for advanced searching and mining of code and code snippets on GitHub, often used to find sensitive informati…
452152abandoned
minimaxir/facebook-page-post-scraper
A Python script collection that scrapes all posts, reactions, and comments from public Facebook Pages and open Groups via the Facebook Grap…
102135abandoned
simplecrawler/simplecrawler
simplecrawler is a flexible, event-driven web crawler library for Node.js with a configurable queue system, robots.txt support, and link di…
102134abandoned
brenden/node-webshot
A Node.js library providing a simple API for taking website screenshots by wrapping PhantomJS's WebKit rendering. It supports capturing URL…
322109abandoned
althonos/InstaLooter
InstaLooter is a Python CLI tool that downloads pictures and videos from Instagram profiles without using the official API. It is a re-impl…
232098abandoned
fouber/page-monitor
A Node.js library that uses PhantomJS to render webpages, capture screenshots, and diff DOM changes (added/removed elements, text, and styl…
322092abandoned
mukulhase/WebWhatsapp-Wrapper
A Python library providing an unofficial API for WhatsApp by automating WhatsApp Web through Selenium browser automation. It allows sending…
232077abandoned
yahoo/gryffin
Gryffin is a large-scale web security scanning platform written in Go, built on a publisher-subscriber architecture for horizontal scaling.…
102052abandoned
zythum/mama2
MAMA2 is a browser bookmarklet/plugin project that replaces Flash video players on Chinese video sites (Bilibili, Youku, Tudou, Sohu, etc.)…
322040abandoned
anime-dl/anime-downloader
A Python command-line tool for downloading and streaming anime episodes from various streaming sites and Nyaa. It supports batch episode do…
232002abandoned
trevorlinton/webkit.js
A pure JavaScript port of WebKit's WebCore compiled via Emscripten, capable of rendering HTML5, CSS3, and SVG to WebGL/Canvas contexts in b…
321971abandoned
wkeeling/selenium-wire
Selenium Wire extends Selenium's Python bindings to expose the HTTP/HTTPS requests and responses made by the browser, with APIs to inspect …
101965abandoned
mozilla/fathom
Fathom is a supervised-learning framework for recognizing and classifying parts of web pages, tagging DOM nodes with types and probabilitie…
101963abandoned
gxvv/ex-baiduyunpan
A Tampermonkey userscript that removes Baidu Netdisk's large-file download restrictions and enables batch copying of download links. It wor…
101914abandoned
cycz/jdBuyMask
A Python automation tool that monitored JD.com (Jingdong) for face mask stock during the COVID-19 pandemic and automatically placed purchas…
321849abandoned
hu17889/go_spider
go_spider is a concurrent web crawler framework written in Go, designed for crawling vertical communities with a flexible, modular architec…
231818abandoned
NotJoeMartinez/yt-fts
yt-fts is a Python command line tool that downloads all subtitles from a YouTube channel or playlist via yt-dlp and stores them in a SQLite…
591812abandoned
eth0izzle/bucket-stream
A Python CLI tool that monitors certificate transparency logs via certstream and discovers public Amazon S3 buckets by generating permutati…
361808abandoned
timschneeb/tachiyomi-extensions-archive
A historical archive of removed manga source extensions for the Tachiyomi Android reader app. The repository was DMCA'd and erased, and is …
101797abandoned
node-js-libs/node.io
node.io is a Node.js web scraping and data extraction library originally written in 2010. It is explicitly no longer maintained, with the a…
321792abandoned
bughandler/cnki-downloader
A small desktop tool for searching and downloading academic literature from CNKI (China National Knowledge Infrastructure). Its backend int…
481762abandoned
alex/nyt-2020-election-scraper
A git-scraping application that periodically scrapes the New York Times' 2020 election results JSON API and commits snapshots to the reposi…
321756abandoned
mchristopher/PokemonGo-DesktopMap
An Electron desktop application that bundles the PokemonGo-Map project with an HTML UI and Python dependencies to show a live visualization…
101747abandoned
metafates/mangal
Mangal is a cross-platform CLI/TUI manga downloader written in Go, with built-in sources (Mangadex, Manganelo, Manganato, Mangapill), exten…
101745abandoned
sorenlouv/fb-sleep-stats
A Node.js proof-of-concept tool that polls Facebook's online/offline status to infer and visualize friends' sleep patterns. It consists of …
321707abandoned
MShawon/YouTube-Viewer
A Python-based multithreaded YouTube view bot that uses Selenium-driven browser sessions with free, premium, and rotating proxies to inflat…
231663abandoned
cool2528/baiduCDP
BaiduCDP is a C++ Windows application for high-speed downloading from Baidu Netdisk (Baidu Cloud Drive). It analyzes Baidu Netdisk's web AP…
101661abandoned
tmort/Socialite
Socialite is a small vanilla JavaScript library for lazily loading social sharing widgets (Facebook, Twitter, Google+, LinkedIn, Pinterest,…
101660abandoned
sc1341/InstagramOSINT
A Python CLI tool and importable module that scrapes publicly available profile information from Instagram accounts, such as follower count…
101651abandoned
ZFC-Digital/puppeteer-real-browser
A Node.js library that wraps Puppeteer with a real-browser profile to bypass bot detection systems like Cloudflare and Turnstile captchas. …
351642abandoned
junhoyeo/threads-api
An unofficial, reverse-engineered Node.js/TypeScript client for Meta's Threads social network, built shortly after Threads launched, with a…
101621abandoned
erma0/douyin
A Python crawler for Douyin (Chinese TikTok) that collected public data such as account profiles, likes, favorites, music, hashtags, search…
761602abandoned
chinoogawa/fbht
A Python 2 command-line tool for interacting with and scraping Facebook accounts, including graph-based analysis of social connections. It …
231591abandoned
fossasia/loklak_wok_android
Loklak Wok is an Android app that acts as a harvesting peer for the loklak_server, collecting social media messages (tweets) and pushing th…
101558abandoned
fossasia/searss
A small Python CLI tool that scrapes search results from engines like Google, Bing, DuckDuckGo, and Ask.com and converts them into RSS feed…
101534abandoned
GravityLabs/goose
Goose is a Scala library (originally Java) that extracts the main body text, metadata, publish date, embedded videos, and top image from ne…
101526abandoned
fossasia/loklak-webtweets
A small client-side web app that fetches and displays tweets via the loklak API, embeddable on any website with configurable query, count, …
101525abandoned
fossasia/loklak_tweetheatmap
A web application that displays a geographic heat map of tweets matching a search query, built with Angular.js and OpenLayers 3 using the L…
321518abandoned
benitoro/stockholm
A Python framework that crawls Shanghai and Shenzhen A-share stock market data from Yahoo YQL and Sina Finance, and tests stock-picking str…
321517abandoned
githubwing/GankClient-Kotlin
An Android client app for gank.io (a Chinese tech content aggregator) written in Kotlin with Material Design. It serves as a reference impl…
321515abandoned
yhat/scrape
A Go library providing a higher-level interface over golang.org/x/net/html for web scraping. It offers generic tree traversal helpers like …
101515abandoned
FOSSASIA-Web/timeline.api.fossasia.net
A jQuery plugin that embeds a FOSSASIA community event timeline into any website, fetching events from the FOSSASIA/Freifunk Calendar API o…
321514abandoned
isanchop/stuhack
A Chrome extension that attempted to unlock Studocu premium features such as removing premium banners, bypassing document blur, and enablin…
101513abandoned
fossasia/wp-tweet-legacy
A WordPress widget plugin that displays Twitter feeds in widget-ready areas, parsing @usernames, #hashtags, and URLs into links. It support…
321504abandoned
qinxuye/cola
Cola is a high-level distributed crawling framework in Python for scraping pages and extracting structured data from websites. The same cra…
101499abandoned
fossasia/loklak_heatmapper
A plain HTML/JavaScript web application that visualizes recent tweet activity as a heatmap on a world map. It fetches tweets via the loklak…
101497abandoned
artyshko/smd
A Python application that downloads high-quality music from Spotify playlists, albums, and songs by fetching audio from sources like YouTub…
231486abandoned
johntitus/node-horseman
A Node.js library providing a chainable, Promise-based API for controlling the PhantomJS headless browser, supporting page navigation, form…
321485abandoned
yucccc/vue-mall
A full-stack e-commerce demo mall (clone of the Smartisan/锤子 online store) built with Vue 2 + Vuex on the frontend and Node.js + MongoDB fo…
321482abandoned
ring04h/wydomain
wydomain is a Python command-line tool for discovering subdomains of a target domain. It combines dictionary-based DNS bruteforcing with qu…
321479abandoned
fossasia/loklak_scraper_js
A collection of JavaScript scrapers for the loklak project that extract data from websites like Twitter and Quora and output JSON resemblin…
101473abandoned
fossasia/loklak-timeline-plugin
A lightweight JavaScript plugin for embedding loklak search timelines into web pages via a simple HTML tag with data attributes. It renders…
101457abandoned
keenwon/antcolony
AntColony is a Node.js-based BitTorrent DHT network crawler that collects active infohashes, downloads and parses torrent files, and stores…
321456abandoned
Entromorgan/Autoticket
A Python/Selenium automation tool that automatically snipes and purchases tickets on Damai.cn (China's major ticketing site) by driving Chr…
101450abandoned

← prev page 15 / 20 next →