Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
GramAddict/bot
GramAddict is a free, open-source Instagram bot that automates liking, following, commenting, and messaging on any Android device or emulat…
261632active
vigorX777/ai-daily-digest
A zero-dependency TypeScript CLI tool (runnable as an OpenCode skill) that fetches the latest posts from 90 top tech blogs recommended by A…
451630active
lit26/finvizfinance
A Python library that scrapes and downloads financial data from the FinViz website, returning stock fundamentals, technicals, charts, news,…
791623active
vladko312/SSTImap
SSTImap is a Python-based penetration testing tool that automatically detects and exploits Server-Side Template Injection (SSTI) and code i…
881621active
michenriksen/aquatone
Aquatone is a Go CLI tool for visual inspection of websites across many hosts, taking screenshots via headless Chrome/Chromium and generati…
105961maintenance
wu529778790/panhub.shenzjd.com
PanHub is a self-hostable netdisk search aggregator that combines results from Quark, Aliyun Drive, Baidu Netdisk, 115, Thunder and 80+ Tel…
621617active
hartator/wayback-machine-downloader
A Ruby command-line tool that downloads an entire website from the Internet Archive Wayback Machine, restoring original files and directory…
235930maintenance
qsniyg/maxurl
Image Max URL is a userscript and browser extension that finds larger or original versions of images and videos by rewriting URL patterns, …
771611active
ka-pi-ba-la/AIbijia
AIbijia is a price-comparison website and community resource that aggregates AI subscription and token prices across multiple platforms (Ch…
571611active
minkphp/Mink
Mink is a PHP library providing an abstraction layer over web browser emulators and drivers (like Goutte, Selenium) for controlling browser…
531611active
ArchiveTeam/grab-site
grab-site is a preconfigured web crawler for archiving websites, producing WARC files via a fork of wpull. It includes a dashboard for moni…
431607active
matthewmueller/x-ray
x-ray is a Node.js web scraping library that lets you define flexible schemas to structure data from any website using jQuery-like selector…
665907maintenance
Altimis/Scweet
Scweet is a Python library and CLI for scraping tweets, profile timelines, followers, following lists, and user profiles from Twitter/X wit…
781604active
Xatta-Trone/medium-parser-extension
A browser extension for Chrome, Edge, and Firefox that lets users read member-only articles on medium.com and Medium-based sites (e.g. Towa…
591589active
lanyeeee/jmcomic-downloader
A multi-threaded GUI downloader for the 18comic.vip (jmcomic) manga site, built with Tauri (Rust backend, Vue frontend). It supports search…
891587active
Autumn-27/ScopeSentry
ScopeSentry is a self-hosted attack surface and asset mapping platform that combines subdomain enumeration, port scanning, fingerprinting, …
881587active
s0md3v/uro
uro is a Python CLI tool that declutters URL lists for crawling and security testing without making any HTTP requests. It removes duplicate…
291587stable
xnl-h4ck3r/xnLinkFinder
xnLinkFinder is a Python CLI tool that discovers endpoints, potential parameters, target-specific wordlists, and secrets for a given target…
781585active
m3n0sd0n4ld/GooFuzz
GooFuzz is a Bash-based CLI tool that performs fuzzing-style reconnaissance using advanced Google searches (Google Dorking) via the Google …
571585active
LeyckerS/moondownloader
Moon Downloader is a bulk file downloader for the datanodes.to and fuckingfast.co file hosts, extracting direct links either via a real Chr…
881584active
actionbook/actionbook
Actionbook is an AI agent that lives in a Chrome side panel and operates websites on your behalf, reading data behind logins and paywalls a…
741584active
AlisamTechnology/ATSCAN
ATSCAN is a Perl-based command-line scanner for mass dork searching and vulnerability exploitation. It combines search engine dorking with …
231583active
m8sec/CrossLinked
CrossLinked is a Python CLI tool that enumerates LinkedIn employee names for an organization by scraping search engine results, without nee…
231582active
agentrhq/webcmd
Webcmd is a self-learning browser infrastructure CLI for AI agents that compiles knowledge of websites into deterministic commands, cutting…
801579active
postlight/parser
Postlight Parser is a JavaScript library that extracts meaningful content—article text, titles, authors, dates, lead images, and excerpts—f…
235786maintenance
Leon406/SubCrawler
A Kotlin-based tool that automatically crawls and health-checks (via Google ping) free public proxy nodes for protocols like V2Ray, Shadows…
771561active
OWASP/QRLJacking
QRLJacking is an OWASP project documenting and exploiting the Quick Response Code Login Jacking attack vector, which hijacks user sessions …
481559active
ion-design/ditto.site
ditto.site is a hosted capture-to-code service that turns a public URL into a self-contained TypeScript app, emitting a deterministic Next.…
571558active
ttttmr/Wechat2RSS
Wechat2RSS is a service and self-hostable tool that converts WeChat official account (公众号) articles into RSS feeds, aiming for updates with…
751557active
tafia/quick-xml
quick-xml is a high-performance XML pull reader and writer library for Rust, featuring near zero-copy parsing with Cow types and buffer reu…
951556stable
webrecorder/archiveweb.page
ArchiveWeb.page is a high-fidelity web archiving tool that runs as a Chrome/Chromium browser extension and as a standalone Electron app. It…
931555active
ulixee/hero
Hero is a headless web browser built specifically for web scraping, powered by Chrome and controlled from NodeJS with a fully compliant DOM…
691554active
littledivy/mimic
mimic is a Python tool that captures traffic from mobile or web apps via mitmproxy, extracts authentication material, and uses AI to genera…
551550active
Rhizome-Conifer/conifer
Conifer is an open-source web archiving platform for capturing, replaying, and sharing collections of archived web pages through a user-fri…
741549active
skernelx/tavily-key-generator
A Python toolkit that automates signup flows for Tavily, Firecrawl, and Exa using real browser automation (Playwright/Camoufox), Turnstile …
481549active
yujiosaka/headless-chrome-crawler
A Node.js library providing a distributed web crawler powered by Headless Chrome via Puppeteer. It can crawl JavaScript-rendered (SPA) webs…
235635maintenance
ecmadao/hacknical
Hacknical is a web application that analyzes a GitHub user's data (contributions, commits, languages, repos) and helps generate a better de…
491542active
hyperbrowserai/HyperAgent
HyperAgent is a TypeScript library and CLI that adds LLM-powered natural language commands to Playwright for browser automation. It support…
521540active
DialmasterOrg/Youtarr
Youtarr is a self-hosted web application that automatically downloads YouTube channel and playlist content, organizes it with metadata for …
921539active
oxylabs/ai-map-py
AI-Map is a Python SDK client for Oxylabs AI Studio's AI-powered website mapping service, which discovers and extracts relevant URLs from a…
501537active
FinanceData/FinanceDataReader
FinanceDataReader is a Python library and CLI tool for reading financial data such as stock listings, stock prices, indexes, exchange rates…
691535active
cross-seed/cross-seed
cross-seed is a Node.js application that automatically finds and downloads torrents matching your existing torrent library across multiple …
921534active
LifeActor/ykdl
YouKuDownLoader (ykdl) is a Python command-line video downloader focused on China mainland video sites, forked from you-get with restructur…
471533active
xnl-h4ck3r/GAP-Burp-Extension
GAP is a Burp Suite extension written in Python (Jython) that extracts potential endpoints, parameters, and links from Burp's site map, pro…
661530active
rbren/rss-parser
rss-parser is a lightweight JavaScript library that converts RSS and Atom XML feeds into JavaScript objects. It works in both Node.js and t…
771527active
justfoolingaround/animdl
animdl is a lightweight Python CLI tool that scrapes, streams, and downloads anime episodes from supported providers. It supports quality s…
321522active
tidyverse/rvest
rvest is an R package from the tidyverse for scraping (harvesting) data from web pages, inspired by Beautiful Soup and RoboBrowser. It prov…
431520active
SpiderClub/haipproxy
A high-availability distributed IP proxy pool built with Scrapy and Redis that scrapes free proxies from the internet, validates them, and …
235523maintenance
jasperan/whatsapp-osint
A Python CLI tool that uses Selenium to track when WhatsApp contacts go online/offline, logging presence sessions to SQLite. It exports dat…
761511active
Pickle-Pixel/ApplyPilot
ApplyPilot is an open-source AI agent that autonomously applies to jobs on your behalf across any job site or application form. It runs a 6…
601508active
oxylabs/browser-agent-py
A Python SDK for Oxylabs AI Studio's Browser Agent, a cloud service that automates real-user browsing tasks (clicking, typing, scrolling, s…
511505active
manojVivek/medium-unlimited
A browser extension for Chrome and Firefox that unlocks Medium.com membership-only articles by bypassing the paywall. It supports medium.co…
105467maintenance
JustAnotherArchivist/snscrape
snscrape is a Python-based scraper for social networking services that extracts posts, profiles, hashtags, and search results from platform…
325444maintenance
rafatosta/zapzap
ZapZap is an unofficial WhatsApp Web desktop client built with Python, PyQt6, and QtWebEngine that wraps web.whatsapp.com in a native deskt…
951495active
Vincentqyw/cv-arxiv-daily
An automated daily digest of computer vision and robotics arXiv papers (SLAM, SFM, visual localization, keypoint detection, image matching,…
771494active
SilentDemonSD/WZML-X
WZML-X is a self-hosted Telegram mirror and leech bot written in Python that downloads files from torrents, Mega, Google Drive, direct link…
961493active
jlesage/docker-jdownloader-2
A Docker container packaging JDownloader 2, a download manager, with a browser-accessible graphical interface via web or VNC. It is an unof…
981492active
Tsuk1ko/bilibili-live-chat
A backend-free web app that displays Bilibili live stream danmaku (chat messages) and gifts in a YouTube Live Chat style overlay, primarily…
601486active
exa-labs/company-researcher
An open-source web application by Exa.ai that lets users enter a company URL and instantly gathers comprehensive research about it, includi…
651483active
AI4Finance-Foundation/FinNLP
FinNLP is a Python library for collecting internet-scale financial data from sources like Finnhub, Yahoo Finance, Reuters, and Sina Finance…
311480active
sethblack/python-seo-analyzer
A Python-based SEO analyzer that crawls a website, analyzes its structure, counts body words, and reports technical SEO issues, with option…
651475active
Ademking/MD-This-Page
A browser extension for Chrome and Firefox that converts any webpage into clean, readable Markdown with one click, using Mozilla's Readabil…
611475active
Surfer-Org/Protocol
Surfer Protocol is an open-source framework for exporting personal data from platforms like Gmail, iMessages, Twitter, Notion, and ChatGPT.…
131475active
shuanx/BurpAPIFinder
BurpAPIFinder is a Burp Suite extension written in Java that passively analyzes HTTP traffic (HTML and JS files) to discover hidden API end…
151472active
epsylon/xsser
XSSer is an automatic penetration testing framework for detecting, exploiting, and reporting cross-site scripting (XSS) vulnerabilities in …
821461active
momosecurity/FindSomething
FindSomething is a passive browser extension for Chrome and Firefox that extracts potentially sensitive information (like emails, API keys,…
321458active
mylar3/mylar3
Mylar3 is a Python-based automated comic book (cbr/cbz) downloader that monitors a watchlist of series and grabs new issues via NZB indexer…
641455active
roach-php/core
Roach is a complete web scraping and crawling toolkit for PHP, heavily inspired by Python's Scrapy. It lets developers define spiders that …
451455active
tychxn/jd-assistant
A JD.com (Jingdong) purchase assistant written in Python that automates login via QR code, product stock/price queries, cart management, an…
325259maintenance
submato/xhscrawl
A Python-based reverse-engineering toolkit for Xiaohongshu (XHS) web APIs, focusing on generating the encrypted x-s signature parameter via…
721452active
tinyfish-io/agentql
AgentQL is a suite of tools for extracting structured data and automating workflows on live websites using an AI-powered natural language q…
701451active
gpodder/gpodder
gPodder is a free, open-source podcast client and media aggregator that lets users subscribe to, download, and manage podcast feeds. It has…
661451stable
eliasdabbas/advertools
advertools is a Python library of productivity and analysis tools for online/digital marketing, built as a set of independent, composable f…
731446active
orangecoding/fredy
Fredy is a self-hosted Node.js application that continuously scrapes European real estate portals like ImmoScout24, Immowelt, Kleinanzeigen…
951444active
6551Team/opentwitter-mcp
A Python MCP server that exposes Twitter/X data (user profiles, tweet search, follower events, deleted tweets, KOL tracking) to AI assistan…
591444active
denandz/sourcemapper
Sourcemapper is a Go CLI tool that parses JavaScript sourcemap (.map) files, whether from URLs or local directories, and reconstructs the o…
751442stable
txperl/PixivBiu
PixivBiu is a self-hosted Pixiv client written in Go that provides a browser-based interface for searching, filtering, browsing, and downlo…
871440active
scrapy-plugins/scrapy-playwright
A Scrapy download handler that uses Playwright for Python to fetch pages, enabling scraping of JavaScript-rendered sites while keeping the …
861439active
nickclyde/duckduckgo-mcp-server
A Model Context Protocol (MCP) server that exposes DuckDuckGo web search to LLM clients like Claude Desktop and Claude Code. It also fetche…
831439active
asz798838958/aBaiFreeGPT
A self-hosted account registration and lifecycle management platform that automates bulk account signup, email verification, TOTP 2FA bindi…
581438active
ShilongLee/Crawler
A self-hostable crawler API server that exposes HTTP endpoints for scraping public data from Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo…
531435active
apinanaivot/IKEA-3D-Model-Download-Button
A Tampermonkey userscript that adds a 'Download 3D' button to IKEA product pages, letting users save the product's 3D model as a .GLB file.…
531435active
drawrowfly/tiktok-scraper
A TypeScript library and CLI tool that scrapes TikTok metadata from user, hashtag, trend, and music pages and downloads video posts without…
235173maintenance
sw33tLie/bbscope
bbscope is a Go CLI tool that fetches, stores, and manages bug bounty program scopes from HackerOne, Bugcrowd, Intigriti, YesWeHack, and Im…
731432active
wd210010/only_for_happly
A collection of Python automation scripts designed to run on the Qinglong (青龙) panel, covering daily check-ins for services like Baidu Tieb…
741431active
dteviot/WebToEpub
A Chrome and Firefox browser extension that converts web novels and other web pages into EPUB ebooks for offline reading. It supports many …
941430active
kepano/clipper-templates
A collection of templates for the Obsidian Web Clipper browser extension, covering generic schemas (recipes, products) and specific sites l…
551427active
zhaoolee/garss
Garss (嘎!RSS) is a self-hosted RSS aggregation and reading system that uses GitHub Actions to collect hundreds of RSS feeds and render them…
821426active
lanlinju/Animius
Animius is a clean, minimalist Android app for watching anime, built with Jetpack Compose and Kotlin. It supports danmaku (bullet comments)…
821425active
rebrowser/rebrowser-patches
A collection of source-code patches for Puppeteer and Playwright that fix automation leaks and help avoid bot detection systems like Cloudf…
341424active
stereobooster/react-snap
A zero-configuration, framework-agnostic static prerendering tool for single-page applications. It uses Headless Chrome (via Puppeteer) to …
625116maintenance
lorenzodifuccia/safaribooks
A Python CLI tool that downloads books from O'Reilly Learning (Safari Books Online) and generates EPUB files from them. It requires a valid…
515100maintenance
openwpm/OpenWPM
OpenWPM is a web privacy measurement framework built on Firefox with Selenium automation, designed to collect data from thousands to millio…
951417active
danny0838/content-farm-terminator
Content Farm Terminator is a cross-platform browser extension that identifies content farms by marking hyperlinks pointing to them and bloc…
771416active
debridmediamanager/debrid-media-manager
Debrid Media Manager (DMM) is a free, open-source web app for curating an unlimited-size movie and TV show library on debrid services like …
751416active
hxh19950701/WebViewTvLive
An Android TV live streaming app that loads official broadcaster web pages in a Tencent X5 WebView, auto-fullscreens the video element, and…
875094maintenance
facundoolano/app-store-scraper
A Node.js library for scraping application data from the iTunes and Mac App Store. It provides methods to retrieve app details, search resu…
471414active
sec-edgar/sec-edgar
A Python library and CLI for downloading company periodic reports, filings, and forms from the SEC's EDGAR database. It supports fetching f…
481413active
cdpdriver/zendriver
Zendriver is an async-first Python web scraping and browser automation framework built on the Chrome Devtools Protocol, forked from nodrive…
881409active
browserwing/browserwing
BrowserWing is an open-source browser automation platform written in Go with a React dashboard that exposes browser control via MCP command…
711408active

← prev page 8 / 20 next →