Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
hengliyin/cdfang-spider
A full-stack web application that scrapes Chengdu Housing Association lottery housing data and presents it through interactive charts and s…
531221active
warifp/FacebookToolkit
A PHP command-line toolkit for retrieving Facebook account data via the Graph API, including access tokens, friend IDs, emails, and names, …
231221active
cnbattle/douyin
A Go-based crawler that scrapes Douyin (TikTok China) recommendation and search page video lists by controlling the mobile app on a real de…
611219active
Silent1566/OmniBox-Spider
A collection of spider (scraper) sources and interfaces for the OmniBox media application, aggregated from publicly available internet info…
581213active
Sak32009/GetDataFromSteam-SteamDB
A userscript that extracts game and DLC data from Steam store pages and SteamDB. It runs via userscript managers like Tampermonkey or Viole…
761212active
taranis-ai/taranis-ai
Taranis AI is a self-hosted open-source OSINT platform that collects news articles from web sources and uses NLP/AI to enrich, cluster, and…
951207active
lexmount/moli
Moli is a lightweight, fast headless browser built in Rust (on Servo technology) designed for AI agents to fetch, render, and extract web p…
791206active
browserable/browserable
Browserable is an open-source, self-hostable JavaScript library for building AI browser agents that navigate sites, fill out forms, click b…
371206active
CERT-Polska/Artemis
Artemis is a modular, open-source vulnerability scanner developed by CERT Polska that checks website security at scale. It automatically ge…
741204active
gityuanbao/share
A personal open-source repository whose main component is akshare_collector, a Python tool built on AKShare that collects Chinese financial…
641201active
scrapinghub/splash
Splash is a lightweight, scriptable headless browser exposed as a service with an HTTP API, implemented in Python 3 using Twisted and Qt5. …
324187maintenance
ycdxsb/PocOrExp_in_Github
A Python CLI tool that automatically aggregates proof-of-concept (POC) and exploit (EXP) code from GitHub by CVE ID, using CVE information …
771198active
RyensX/MediaBox
MediaBox is an Android 'universal media container' app that aggregates video, manga, and other media through a WeChat-mini-program-like plu…
231198active
internetwache/GitTools
GitTools is a collection of three shell/Python scripts for finding and exploiting websites that publicly expose their .git directory. It in…
644178maintenance
niudaii/zpscan
zpscan is a Go-based command-line information gathering and reconnaissance tool for security assessments. It bundles subdomain enumeration,…
231196active
instantbox/instantbox
instantbox is a self-hosted service that spins up temporary, clean Linux containers (Ubuntu, CentOS, Arch, Debian, Fedora, Alpine) with ins…
324176maintenance
teaSummer/MCiSEE
MCiSEE is a web application that aggregates and presents Minecraft resources (software, tools, websites) collected from the internet in an …
691194active
CarrotRub/Fit-Launcher
Fit Launcher is a desktop game launcher built with Rust, Tauri, and SolidJS for downloading and playing games from FitGirl Repacks. It uses…
771193active
constverum/ProxyBroker
ProxyBroker is an asynchronous Python tool that finds public HTTP(S) and SOCKS4/5 proxies from ~50 sources and concurrently checks their ty…
324159maintenance
jayus0821/swagger-hack
A Python CLI tool that automatically crawls all endpoints exposed by leaked Swagger/OpenAPI documentation and sends configured test request…
571189active
xenova/chat-downloader
Chat Downloader is a Python tool and library for retrieving chat messages from livestreams, videos, clips, and past broadcasts on platforms…
451189active
intoli/user-agents
A JavaScript/TypeScript npm package for generating random user agents weighted by real-world market share, with daily-updated data. It also…
771188active
kmille/deezer-downloader
A self-hosted Python service with a simple web frontend for downloading songs, albums, and playlists from Deezer (and via yt-dlp), with ID3…
741186active
neon-mmd/websurfx
Websurfx is an open-source meta search engine written in Rust that aggregates results from multiple search engines while respecting user pr…
861184active
LainsNL/OutlookRegister
A Python automation tool that mass-registers Outlook/Hotmail email accounts using browser automation (Playwright or Patchright) with simula…
621184active
gooseworks-ai/goose-skills
A library of 200+ AI agent skills and data APIs for growth and go-to-market work, installable into coding agents like Claude Code, Cursor, …
581184active
goodreasonai/ScrapeServ
ScrapeServ is a self-hosted API service that accepts a URL and returns the website's data along with browser screenshots, using Playwright …
251181active
shobrook/rebound
Rebound is a command-line tool that runs your file and, when an exception is thrown, instantly fetches related Stack Overflow questions and…
234114maintenance
austin-weeks/miasma
Miasma is a lightweight Rust web server that traps AI web scrapers in an endless pit of poisoned training data and self-referential links. …
821179active
ycngmn/Nobook
Nobook is a lightweight, ad-free Android client for browsing Facebook, built with Kotlin and Jetpack Compose on top of a WebView. It blocks…
101179active
rchipka/node-osmosis
Osmosis is an HTML/XML parser and web scraper library for Node.js built on native libxml C bindings. It offers a chainable, promise-like in…
324107maintenance
grangier/python-goose
Python-Goose is a Python library that extracts the main body text, metadata, top image, and embedded videos from news article web pages. It…
644106maintenance
scdl-org/scdl
scdl is a Python command-line tool for downloading music, playlists, likes, and reposts from SoundCloud, with automatic ID3 metadata taggin…
674099maintenance
N0rz3/Phunter
Phunter is a Python CLI OSINT tool that gathers information about phone numbers, including operator, line type, location, reputation, spam …
261175active
apache/groovy-geb
Apache Geb is a Groovy-based browser automation library built on WebDriver, combining jQuery-like content selection with Page Object modell…
761173active
myfanhua/turb-gpt-free-register
A Python tool that bulk-registers ChatGPT/OpenAI accounts using pure-protocol requests or anti-fingerprint browser automation (RoxyBrowser,…
581172active
Evolution0/bandcamp-dl
A Python command-line tool for downloading albums and tracks from bandcamp.com, with options for filename templating, album art embedding, …
651171active
online-judge-tools/oj
A command-line tool that automates solving problems on online judges like AtCoder, Codeforces, and HackerRank. It downloads sample and syst…
231169active
rverton/webanalyze
webanalyze is a Go port of Wappalyzer that detects the technologies used on websites, built for performant mass scanning of large host list…
711168active
aooiuu/any-reader
Any-Reader is an open-source, cross-platform content aggregation tool that unifies novels, manga, video, and audio from user-defined source…
731167active
xeco23/WasIstLos
WasIstLos is an unofficial WhatsApp desktop client for Linux, written in C++ using gtkmm and WebKitGTK to wrap WhatsApp Web. It adds deskto…
101166active
asdfghj1237890/WebVideo2NAS
A self-hosted pipeline consisting of a Chrome extension and a Dockerized FastAPI backend that detects HLS (M3U8), DASH (MPD), MP4, and MOV …
801161active
CMHopeSunshine/LittlePaimon
LittlePaimon is a multifunctional Genshin Impact chatbot built on the NoneBot2 framework, supporting the OneBot protocol for QQ. It provide…
271160active
zu1k/proxypool
A Go service that automatically crawls proxy nodes (ss, ssr, vmess, trojan) from Telegram channels, subscription URLs, and the public inter…
234027maintenance
0x727/ShuiZe_0x727
ShuiZe_0x727 is a Python-based automated information gathering (reconnaissance) tool for red team operators. Given a root domain, C-segment…
234019maintenance
fanpei91/torsniff
torsniff is a Go CLI tool that sniffs torrent metadata from the BitTorrent network by participating in the DHT and connecting to peers to d…
104014maintenance
librariesio/libraries.io
Libraries.io is an open source discovery service that indexes millions of packages across 32 package managers, letting developers search an…
751156active
evyatarmeged/Raccoon
Raccoon is a Python-based offensive security CLI tool for reconnaissance and information gathering. It performs DNS lookups, WHOIS, TLS ana…
674001maintenance
EmilStenstrom/justhtml
JustHTML is a pure Python HTML5 parser with browser-style error recovery, safe-by-default sanitization, CSS selector querying, and serializ…
831151active
agentbay-ai/wuying-agentbay-sdk
Multi-language SDK (Python, TypeScript, Go, Java) for Wuying AgentBay, Alibaba Cloud's cloud sandbox platform built for AI agents. It lets …
761148active
filipedeschamps/rss-feed-emitter
A Node.js library that aggregates RSS and Atom news feeds and emits events for every new item published. It automatically manages feed hist…
721147stable
yunginnanet/HellPot
HellPot is a cross-platform HTTP honeypot that punishes bots ignoring robots.txt by streaming an infinite Markov-chain-generated page of ps…
491147active
random-robbie/My-Shodan-Scripts
A collection of Python 3 scripts for querying the Shodan search engine to find exposed devices and services on the internet. It bundles man…
601146active
0xsha/CloudBrute
CloudBrute is a Go CLI tool that enumerates a company's infrastructure, files, and applications across major cloud providers (Amazon, Googl…
271145active
datawhores/OF-Scraper
OF-Scraper is a command-line tool for downloading media from OnlyFans and performing bulk actions like liking or unliking posts. It is a re…
801143active
nicolomantini/LinkedIn-Easy-Apply-Bot
A Python bot that automates applying to jobs on LinkedIn using the Easy Apply feature. It reads search preferences, resume uploads, and bla…
571143active
Tucsky/aggr
AGGR (SignificantTrade) is a Vue.js web application that aggregates live cryptocurrency trades from many exchanges (Binance, Coinbase, BitM…
591142active
AndyTheFactory/newspaper4k
Newspaper4k is a Python library and CLI for scraping and curating news articles, extracting text, titles, authors, publish dates, and metad…
841140active
kevthehermit/PasteHunter
PasteHunter is a Python 3 application that queries public pastebin-style sites (pastebin.com, GitHub gists, slexy, stackexchange, etc.) and…
501137active
ChineseSubFinder/ChineseSubFinder
A self-hosted Go application that automatically downloads Chinese subtitles for movies and TV shows from subtitle websites. It integrates w…
233932maintenance
fwonggh/Bthub
Bthub is a magnet link and torrent search engine, and this repository serves as its official address release page listing current and backu…
761133active
hasanfirnas/symbiote
Symbiote is a Python-based social engineering tool that generates a phishing page to trick a target into granting camera permission, then c…
371132active
auto-novel/auto-novel
AutoNovel is a website application that automatically machine-translates light novels (web novels, bunko novels, and local files) using LLM…
771130active
rebane2001/xikipedia
Xikipedia is a web application that presents Wikipedia content as a social media feed, using a simple non-ML algorithm that learns user eng…
531129active
codelibs/fess
Fess is an open-source, self-hosted enterprise search server built on OpenSearch with a browser-based admin UI, built-in crawlers for web s…
951128active
johnwmillr/LyricsGenius
A Python client library that wraps the Genius.com API and scrapes song lyrics, artist, and album metadata. It provides simple search and do…
691121active
QIN2DIM/epic-awesome-gamer
A Python application that automatically claims weekly free games and monthly content from the Epic Games Store. It uses browser automation …
561121active
webrecorder/browsertrix-crawler
Browsertrix Crawler is a standalone browser-based high-fidelity web crawling system that runs in a single Docker container. It uses Puppete…
991120active
Haleydu/Cimoc
Cimoc is an open-source Android manga/comic reader app that aggregates multiple online comic sources. It supports page and scroll reading m…
233855maintenance
r3nt0n/bopscrk
bopscrk is a Python CLI tool that generates smart, targeted wordlists for password cracking, combining user-provided words with transformat…
231117active
Doriandarko/make-it-heavy
A Python framework that emulates Grok Heavy-style deep analysis by orchestrating multiple specialized AI agents in parallel via OpenRouter'…
321116active
shanmiteko/LotteryAutoScript
A Node.js automation script that monitors Bilibili dynamic posts for lottery/giveaway draws and automatically participates by liking, comme…
891115active
platonai/Browser4
Browser4 is an AI-native browser engine built in Kotlin for autonomous agents, intelligent data extraction, and large-scale web automation.…
1001114active
elixir-crawly/crawly
Crawly is a high-level web crawling and scraping framework for Elixir, modeled after Scrapy, where developers define spiders that fetch pag…
371114active
nottelabs/reverse-api-engineer
Reverse API Engineer is a Python CLI tool that captures browser network traffic (HAR) from a website and uses a configured AI model to gene…
841113active
gautamkrishnar/socli
SoCLI is a Python command-line client for Stack Overflow that lets developers search and browse questions and answers without leaving the t…
561111active
bellingcat/auto-archiver
A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supp…
921109active
mrkrsl/web-search-mcp
A locally hosted MCP (Model Context Protocol) server written in TypeScript that gives local LLMs web search capabilities without requiring …
361109active
gxtrobot/bustag
Bustag is a self-hosted web application that periodically scrapes new media title (番号) listings, lets users tag items as liked or disliked,…
233814maintenance
leovan/SciHubEVA
SciHubEVA is a cross-platform GUI application for searching and downloading papers from Sci-Hub, built with Python and Qt (PySide6/QML). It…
741107active
Komet/MediaElch
MediaElch is a cross-platform desktop media manager for Kodi that scrapes and organizes metadata for movies, TV shows, concerts, and music.…
671102active
vifreefly/kimuraframework
Kimuraframework (Kimurai) is a Ruby web scraping framework with an AI-assisted DSL: an LLM generates XPath selectors from a schema on first…
611102active
GPTaku Plugins
insane-search is a Claude Code plugin that reads public web pages that would otherwise be blocked (403, CAPTCHA, WAF), escalating through p…
601102active
alvarorichard/GoAnime
GoAnime is a terminal-based (TUI) anime browser written in Go that lets users search, stream, and download anime episodes directly in mpv. …
941101active
requireCool/stealth.min.js
A repository that automatically generates and publishes the newest stealth.min.js file every Monday at 6 AM UTC. The file is a drop-in Java…
771099active
iFurySt/RedNote-MCP
A Model Context Protocol (MCP) server that lets AI clients access RedNote (Xiaohongshu) content, including keyword search and note retrieva…
191097active
SteamTracking/SteamTracking
A project that tracks and reverse-engineers Steam and Valve data, including protobuf definitions and changes across Steam services. It moni…
771095active
projectdiscovery/wappalyzergo
A high-performance Go library that ports the Wappalyzer technology detection stack, identifying web technologies from HTTP headers and HTML…
951092active
eshaham/israeli-bank-scrapers
A TypeScript library providing scrapers for all major Israeli banks and credit card companies, published as the npm package israeli-bank-sc…
961091active
Tuhinshubhra/RED_HAWK
RED_HAWK is a PHP-based all-in-one reconnaissance and vulnerability scanning tool for websites. It performs information gathering (whois, D…
323748maintenance
YaoZeyuan/stablog
Stablog (稳部落) is a desktop application that backs up and exports a user's Weibo posts into searchable HTML and PDF ebooks. It logs into the…
661087active
GeneralMills/pytrends
Pytrends is an unofficial Python library providing a pseudo API for Google Trends, enabling automated downloading of trend reports. It wrap…
103730maintenance
JimmXinu/FanFicFare
FanFicFare is a Python tool that downloads stories from over 100 fanfiction and web fiction sites and converts them into EPUB (and HTML) eB…
981085active
sharebook-kr/pykrx
PyKrx is a Python library that scrapes stock and bond market data from the Korea Exchange (KRX) and Naver. It provides APIs for querying ti…
761084active
Tyrrrz/YoutubeExplode
A .NET library providing an abstraction layer over YouTube's internal API to query metadata for videos, playlists, and channels, and to res…
993717maintenance
ruipgil/scraperjs
Scraperjs is a Node.js web scraping library offering two scrapers: a lightweight StaticScraper using cheerio for static HTML, and a Dynamic…
323714maintenance
techtanic/Discounted-Udemy-Course-Enroller
A Python application (with GUI and CLI variants) that scrapes websites for 100% off Udemy course coupons and automatically enrolls the user…
711081active
nasa/apod-api
NASA's open-source microservice that serves the Astronomy Picture of the Day (APOD) API, returning JSON metadata and image links parsed fro…
651080active
viu-media/viu
Viu is a terminal-based anime client that provides a rich TUI for browsing, searching, and managing your AniList library, along with stream…
101079active
m-sec-org/EZ
EZ is a cross-platform vulnerability scanner that combines information gathering, port scanning, service brute-forcing, URL crawling, finge…
241078active

← prev page 10 / 20 next →