Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
jmcarp/robobrowser
RoboBrowser is a Pythonic library for browsing the web without a standalone browser, combining Requests for HTTP sessions with BeautifulSou…
323692maintenance
jae-jae/fetcher-mcp
A Model Context Protocol (MCP) server that fetches web page content using a Playwright headless browser, executing JavaScript to handle dyn…
481076active
linkchecker/linkchecker
LinkChecker is a GPL-licensed Python tool that checks links in web documents or entire websites for broken URLs. It supports recursive mult…
651075active
1061700625/WeChat_Article
A PyQt5 desktop application that crawls and downloads all articles from a specified WeChat official account. It uses Selenium to log in and…
621075active
scrapfly/scrapfly-scrapers
A collection of educational Python web scraping scripts for over 40 popular domains such as Amazon, AliExpress, BestBuy, and Twitter, built…
741074active
xuejianxianzun/PixivFanboxDownloader
A Chrome browser extension for batch downloading files from Pixiv Fanbox posts. It supports file type filtering, custom filenames, multiple…
981073active
soxoj/socid-extractor
socid_extractor is a Python library and CLI that extracts structured account metadata and stable internal identifiers (usernames, UIDs, GAI…
921073active
Junyi-99/ChatGPT-API-Scanner
A Python CLI tool that scans GitHub for publicly leaked OpenAI API keys using Selenium browser automation. It is intended for security rese…
581073active
tomnomnom/assetfinder
A Go command-line tool that discovers domains and subdomains potentially related to a given domain by querying multiple passive sources lik…
233666maintenance
borisbabic/browser_cookie3
A Python library (fork of browsercookie) that loads cookies from installed web browsers like Chrome, Firefox, Edge, Safari, and others into…
231070active
jikan-me/jikan
Jikan is an unofficial PHP library and REST API for MyAnimeList.net that scrapes the website to provide data the official API lacks. It let…
641066active
QingJ01/123pan_unlock
A Tampermonkey userscript that unlocks download restrictions on the 123pan (123云盘) cloud storage service, including bypassing the 1GB downl…
101062active
jez500/pricebuddy
PriceBuddy is a self-hostable web application that tracks product prices and availability across online stores, keeping price history and n…
861060active
unitedstates/congress
A community-run Python toolkit that collects and converts official U.S. Congress data—bills, amendments, roll call votes, nominations, and …
521060active
cantino/selectorgadget
SelectorGadget is an open-source bookmarklet and Chrome extension that generates CSS selectors for page elements through point-and-click se…
691059active
spider-ios/autox-release
AutoX is a desktop social media operations tool that automates one-click publishing of videos to multiple platforms such as Douyin, TikTok,…
651059active
meowcateatrat/elephant
Elephant is an add-on for Free Download Manager that adds support for downloading videos from various websites, powered by YT-DLP. It is di…
981058active
chainreactors/spray
Spray is a high-performance HTTP directory fuzzing and content discovery tool written in Go, positioned as a next-generation alternative to…
931058active
Cinvin/myuserscripts
A Tampermonkey userscript for NetEase Cloud Music's web player that adds song downloading, cloud-disk transfer, fast cloud-disk upload, and…
761058active
Rain120/qq-music-api
A QQ Music API service built with Koa2 and TypeScript that proxies web-side QQ Music endpoints for songs, artists, playlists, and rankings.…
761056active
lennybase/browsernode
Browsernode is a TypeScript implementation of Browser-use that lets LLM-powered AI agents control a web browser via Playwright. It provides…
341055active
Casvt/Kapowarr
Kapowarr is a self-hosted web application for building and managing a digital comic book library, designed to fit into the *arr suite of me…
871054active
kort0881/telegram-proxy-collector
A Python CLI tool that automatically collects, analyzes, and filters MTProto and SOCKS5 proxies for Telegram. It decodes proxy secrets to d…
791054active
Zarcolio/sitedorks
A Python CLI tool that runs Google dork-style searches across multiple search engines (Google, Bing, DuckDuckGo, Yandex, Yahoo, Ecosia, Bra…
761053active
Cloxl/xhshow
A pure-algorithm Python library that generates Xiaohongshu (XHS/RedNote) request signature headers such as x-s, x-s-common, x-t, and x-rap-…
761052active
dsclca12/auto_reg
Any Auto Register is a self-hosted multi-platform account automatic registration and management system built with Python and Node.js. It in…
521052active
tombcato/clash-ip-checker
A Python automation tool for Clash proxy users that iterates through proxy nodes, checks each node's IP purity, bot ratio, and IP type via …
441052active
robotshell/magicRecon
MagicRecon is a Bash shell script that automates reconnaissance and vulnerability scanning of target domains, including subdomain enumerati…
231052active
meetDeveloper/freeDictionaryAPI
A free REST API that returns dictionary data for English words, including definitions, phonetics, audio pronunciations, origins, synonyms, …
323590maintenance
Whisparr/Whisparr
Whisparr is an adult movie collection manager for Usenet and BitTorrent users, forked from the Sonarr/Radarr family. It monitors RSS feeds …
1001050active
Henryhaohao/Bilibili_video_download
A Python tool for downloading videos from Bilibili, supporting single-part and multi-part (分P) videos, bangumi episodes, and multiple downl…
323583maintenance
jaimeiniesta/metainspector
MetaInspector is a Ruby gem for web scraping that fetches a given URL and exposes its title, meta description, keywords, links, images, cha…
691049active
lijiejie/GitHack
GitHack is a Python CLI exploit tool that reconstructs a website's source code from an exposed .git folder. It parses the .git/index file, …
323576maintenance
hanFengSan/eHunter
eHunter is a Tampermonkey/userscript that injects a Vue 3-based comic reader UI into supported comic sites (EH/EXHentai, NHentai), offering…
761045active
caolvchong-top/twitter_download
A Python command-line tool that scrapes and downloads images, videos (including GIFs), and text from Twitter/X user timelines. It supports …
751045active
mandatoryprogrammer/thermoptic
Thermoptic is a stealth HTTP proxy that routes requests from any HTTP client (like curl) through a containerized Chrome instance so the tra…
511045active
stay-leave/weibo-public-opinion-analysis
A Python project for Weibo public opinion analysis that combines a web crawler, LDA topic modeling, sentiment analysis, and spatiotemporal …
321045active
turicas/brasil.io
The backend of Brasil.IO, a platform that collects, cleans, and publishes Brazilian public open datasets in accessible formats. It automate…
771044active
tmwgsicp/wechat-download-api
An open-source API service for fetching WeChat official account articles, generating standard RSS 2.0 feeds, and exporting entire account a…
771043active
kelvinBen/AppInfoScanner
A Python-based static information-gathering scanner for mobile apps (Android APK/DEX, iOS IPA/Mach-O) and static web content (HTML, JS, H5)…
233554maintenance
HG-ha/ICP_Query
A self-hosted service (Python and Rust implementations) that queries China MIIT ICP filing records for domains, apps, mini-programs, quick …
951042active
TheAlgorithms/website
The official website for The Algorithms, a static Next.js site that scrapes algorithm implementations from TheAlgorithms GitHub repositorie…
771042active
buffer/thug
Thug is a Python low-interaction honeyclient that mimics the behavior of a web browser to detect and emulate malicious web content. It comp…
871041active
lm-rebooter/NuggetsBooklet
A Node.js script that downloads Juejin (掘金) booklet content for personal study by reusing the user's authenticated browser cookies. It expo…
661041active
badlogic/heissepreise
A self-hostable grocery price search app that daily scrapes product and price data from major Austrian supermarket chains (Billa, Spar, Hof…
601040active
igrigorik/ga-beacon
A small Go service that acts as a collector-as-a-service for Google Analytics via the Measurement Protocol, letting you track page views wi…
323541maintenance
GiantappMan/livewallpaper
Giantapp Livewallpaper is an open-source wallpaper application for Windows 10/11 that supports both dynamic (video/animated) and static wal…
741039active
Anorov/cloudflare-scrape
A Python module (cfscrape) built on Requests that bypasses Cloudflare's JavaScript anti-bot challenge page ('I'm Under Attack Mode') so scr…
233538maintenance
webwhiz-ai/webwhiz
WebWhiz is an open-source, self-hostable application that trains a ChatGPT-powered chatbot on your website data by crawling your pages and …
481037active
TheRook/subbrute
SubBrute is a Python DNS meta-query spider that enumerates subdomains and arbitrary DNS record types by leveraging open resolvers to bypass…
233526maintenance
Diving-Fish/maimaidx-prober
A score tracker (prober) for the arcade rhythm game maimai DX that imports play records via a proxy tool and displays DX Rating and song sc…
911032active
EdgeSecurityTeam/EHole
EHole (棱洞) is a Go-based fingerprint identification tool for red team reconnaissance that pinpoints high-value, easily attackable systems (…
233511maintenance
maxzhang666/OneKeyVip
A multi-function browser userscript (compatible with Tampermonkey and ScriptCat) that bundles VIP video/music parsing, Bilibili cover fetch…
761030active
Wikidepia/InstaFix
InstaFix is a Go web service that serves fixed Instagram image and video embeds for Discord and Telegram by rewriting URLs (e.g., adding 'd…
101029active
carcabot/tiktok-signature
A self-hosted Node.js service that generates valid X-Bogus and X-Gnarly signature tokens for TikTok API requests using a headless browser r…
931025active
wnma3mz/wechat_articles_spider
A Python library for scraping WeChat Official Account articles, including article URLs, reading counts, likes, and comments. It can also do…
233481maintenance
fossology/fossology
FOSSology is an open source license compliance system and toolkit that scans software for licenses, copyrights, and export control data. It…
871024active
agregarr/agregarr
Agregarr is a self-hosted, Docker-based Plex Collections manager that automatically creates and refreshes collections from sources like Tra…
671023active
am-will/codex-skills
A collection of Codex/agent skills written in Shell covering planning, documentation access, prompting, frontend design guidance, Codex too…
571023active
daijro/hrequests
hrequests is a Python HTTP client library that replaces the requests library with browser TLS fingerprint replication, HTTP/2 support, and …
301023active
hrithikkoduri/WebRover
WebRover is an autonomous AI web agent that interprets user input, navigates websites via browser automation, and performs tasks or deep re…
141022active
JosephLai241/URS
URS (Universal Reddit Scraper) is a comprehensive command-line tool written in Python (with Rust components) for scraping and archiving Red…
641020active
owner888/phpspider
phpspider is a PHP web crawling framework that lets developers build scrapers with a simple config array, handling multi-process workers, l…
233462maintenance
mxrch/GitFive
GitFive is a Python-based OSINT CLI tool for investigating GitHub user profiles. It uncovers usernames, name history, email addresses, and …
431019active
wujunwei928/parse-video
A Go library and CLI tool that parses short-video share links from 25+ Chinese platforms (Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo, e…
811018active
Vinyzu/Botright
Botright is a Python browser automation framework built on Playwright that provides undetectable, fingerprint-changing stealth browsing. It…
761017active
vasani-arpit/WBOT
WBOT is a Node.js-based bot for WhatsApp Web that automates message replies using Puppeteer to control a browser. It is configurable via a …
561013active
arabcoders/ytptube
YTPTube is a self-hosted web-based download manager and automation layer for yt-dlp. It combines scheduled tasks, metadata-driven condition…
901012active
davis7dotsh/my-pi-setup
An opinionated configuration and extension setup for the Pi coding agent, adding themes, background terminals, subagents, workflows, and se…
571012active
instagram4j/instagram4j
instagram4j is an object-oriented, reverse-engineered Instagram Private API client library for Java. It lets developers log in, post, messa…
811010active
pea3nut/Pxer
Pxer is a userscript (installed via Tampermonkey) that acts as a crawler for pixiv.net, letting users batch-fetch artworks, collections, an…
271010active
vimcolorschemes/vimcolorschemes
A website and open-source app for browsing and discovering Vim and Neovim colorschemes from GitHub repositories. It tracks thousands of rep…
971007active
wreq
wreq is an ergonomic, privacy-aware HTTP client written in Rust with a Python binding (wreq-python) that provides high-fidelity browser TLS…
901006active
ranahaani/GNews
GNews is a lightweight Python package that queries the Google News RSS feed and returns article results as usable JSON. It supports keyword…
681006active
Kappaemme-git/codex-first-customer-finder-skill
A Codex skill (plugin) that takes a startup URL or product idea and produces an evidence-backed shortlist of potential first customers from…
571006active
JoMingyu/google-play-scraper
A Python library that provides APIs to crawl the Google Play Store for app details, reviews, and other data without any external dependenci…
321006active
ma6254/FictionDown
FictionDown is a Go-based command-line tool for batch downloading and crawling web novels from sites like Qidian and Biquge. It supports mu…
241006active
shy1132/VacuumTube
VacuumTube is an unofficial Electron-based desktop wrapper of YouTube Leanback, the official YouTube interface for consoles and Smart TVs. …
861005active
ki9mu/ARL-plus-docker
A Docker-based fork of ARL (Asset Reconnaissance Lighthouse) v2.6.2 that performs automated asset discovery and vulnerability scanning for …
351005active
TypNull/Tubifarry
Tubifarry is a Lidarr plugin that adds extra music sources to your library management, using Spotify as an indexer and YouTube or Soulseek …
871003active
jackwener/wechat-article-to-markdown
A Python CLI tool that fetches WeChat Official Account articles using anti-detection browser automation (Camoufox) and converts them to cle…
481002active
gwen001/pentest-tools
A collection of small custom security scripts in Bash, Python, and PHP for penetration testing and bug bounty quick tasks, covering DNS enu…
233323maintenance
ping/instagram_private_api
A Python client library for Instagram's private (undocumented) app and web APIs, with no third-party dependencies. It exposes app-only feat…
233299maintenance
s-rah/onionscan
OnionScan is a free and open source Go CLI tool for investigating Tor hidden services (.onion sites) on the Dark Web. It scans sites for op…
233290maintenance
alixaxel/chrome-aws-lambda
A TypeScript library that ships a size-optimized Chromium binary designed to run headless browser automation with Puppeteer (or Playwright)…
323285maintenance
kevinzg/facebook-scraper
A Python library for scraping public Facebook pages, groups, profiles, and posts without requiring an API key. It provides a simple get_pos…
323270maintenance
archivy/archivy
Archivy is a self-hostable personal knowledge repository that combines note-taking, bookmarking with full-page preservation, and a searchab…
233268maintenance
cztomczak/cefpython
CEF Python provides Python bindings for the Chromium Embedded Framework, letting Python applications embed a full Chromium-based browser. I…
593237maintenance
wenbochang888/house
A Java web scraper built with SpringBoot, HttpClient, and JSoup that crawls famous Tianya forum threads about China's housing market and co…
323231maintenance
scrapy-plugins/scrapy-splash
A Scrapy plugin that integrates the Splash headless browser service to enable crawling and scraping of JavaScript-rendered web pages. It pr…
263227maintenance
ferventdesert/Hawk
Hawk is a visual crawler and ETL IDE written in C#/WPF that lets users graphically scrape webpages, clean, transform, and store data withou…
233213maintenance
atlas-comstock/NeteaseCloudMusicFlac
A Python command-line script that downloads lossless FLAC music files to local storage based on a Netease Cloud Music playlist URL. It auto…
323142maintenance
tomnomnom/httprobe
httprobe is a Go command-line tool that takes a list of domains on stdin and probes for working HTTP and HTTPS servers, reporting which res…
233119maintenance
CrawlScript/WebCollector
WebCollector is an open-source Java web crawler framework that provides simple interfaces for building multi-threaded web crawlers quickly.…
623083maintenance
0x0be/yesitsme
A Python CLI script for OSINT investigations that finds Instagram profiles matching a given name, e-mail, or phone number. It scrapes dumpo…
323046maintenance
CodeRayZhang/Movie_Recommend
A full-stack movie recommendation system built on Spark, including a Scrapy crawler, an SSM-based movie website, an admin backend, and a Sp…
323007maintenance
mathieudutour/medium-to-own-blog
A CLI tool that migrates a Medium blog to a self-hosted Gatsby-based blog in minutes. It exports Medium content into Markdown/MDX and scaff…
323003maintenance
jaeles-project/gospider
GoSpider is a fast web spider/crawler written in Go that crawls sites in parallel and extracts URLs from sitemaps, robots.txt, JavaScript f…
232993maintenance
kotartemiy/newscatcher
A Python package that programmatically collects normalized news articles from thousands of news websites, filterable by topic, country, and…
322987maintenance
x0rz/tweets_analyzer
A Python command-line script that scrapes a Twitter profile's tweets and analyzes metadata such as activity patterns, timezone, sources, ge…
232981maintenance

← prev page 11 / 20 next →