Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: crawlers

581 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
MiningCattiva/x-spider
X-Spider is a desktop media downloader for X (Twitter) that fetches images and videos from accounts. It supports media filters, configurabl…
101261abandoned
Zeal-L/BiliBili-Manga-Downloader
A GUI-based Bilibili Manga downloader built with Python and PySide6 that supports keyword search, QR-code login, multi-threaded downloads, …
361242abandoned
syrusakbary/gdom
GDOM is a Python CLI tool that lets you scrape and traverse web page DOM using GraphQL queries, built on the Graphene framework. You write …
321242abandoned
chenjiandongx/mzitu
A Python web crawler that downloads full photo galleries from mzitu.com, collecting over 3,000 image sets. It also performs word-frequency …
321220abandoned
k1995/BaiduyunSpider
A distributed crawler and search engine for Baidu Netdisk (Baidu Cloud) shared links, built with Scrapy and Scrapy-Redis, with a React-base…
231177abandoned
xiaoguyu/wechatDownload
A desktop application built with Electron, TypeScript, and Vue3 for downloading WeChat Official Account (公众号) articles. It captures require…
101170abandoned
shadowmoose/RedditDownloader
Reddit Media Downloader is a Python CLI tool that scans Reddit posts, comments, saved lists, and subreddits to download linked media locall…
101161abandoned
xillwillx/skiptracer
Skiptracer is a Python-based OSINT web scraping framework that aggregates paid and free lookup services to enumerate target information suc…
231146abandoned
ring04h/weakfilescan
A Python-based multi-threaded sensitive information leakage detection tool that crawls a target site, dynamically builds dictionary rules f…
321138abandoned
joshhighet/ransomwatch
A transparent ransomware claim tracker that monitors ransomware groups' leak sites on the dark web and records their victim claims. It aggr…
101116abandoned
bytebuff/JSpider
JSpider is a collection of JavaScript decryption files for website-encrypted request parameters, shared weekly alongside Python snippets th…
321091abandoned
VikParuchuri/apartment-finder
A Python Slack bot that scrapes Craigslist for real-time apartment listings matching configurable criteria (price, neighborhoods, transit p…
101059abandoned
7sDream/zhihu-py3
An unofficial Python 3 API library for Zhihu, the Chinese Q&A site, letting users build objects from Zhihu URLs to fetch questions, answers…
101033abandoned
SpiderClub/smart_login
A Python collection of simulated login implementations for major Chinese websites (Weibo, Zhihu, QQ Zone, JD, Baidu, etc.), using either di…
321008abandoned
puppeteer/puppeteer
Puppeteer is a JavaScript library providing a high-level API to control Chrome or Firefox over the DevTools Protocol or WebDriver BiDi, run…
9595497stable
lightpanda-io/browser
Lightpanda is a headless web browser built from scratch in Zig, designed for AI agents, automation, and web scraping rather than human rend…
9734269active
cheeriojs/cheerio
Cheerio is a fast, flexible JavaScript library for parsing and manipulating HTML and XML with a jQuery-like API. It works in both browser a…
8430468stable
Skyvern-AI/skyvern
Skyvern is an open-source AI browser automation framework that uses LLMs and computer vision to interact with websites, offering a Playwrig…
8822852active
AutomaApp/automa
Automa is a browser extension for Chrome and Firefox that lets users automate browser tasks by visually connecting blocks into workflows. I…
6621586active
chromedp
chromedp is a Go library for driving Chrome and other browsers via the Chrome DevTools Protocol, with no external dependencies. It supports…
8413264active
seleniumbase/SeleniumBase
SeleniumBase is an all-in-one Python browser automation framework for web testing, crawling, and scraping, built on Selenium/WebDriver with…
9512955active
TeamWiseFlow/xiaobei
Xiaobei is an open-source multi-agent system that automates social media marketing and customer acquisition for solo entrepreneurs and smal…
908463active
adithya-s-k/omniparse
OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into…
497815active
go-rod/rod
Rod is a high-level Go library that drives Chrome via the Chrome DevTools Protocol for web automation and scraping. It offers both high-lev…
667077active
lwthiker/curl-impersonate
A special build of curl (and libcurl) whose TLS and HTTP/2 handshakes are identical to real browsers like Chrome, Edge, Safari, and Firefox…
236895active
aidlearning/AidLearning-FrameWork
AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling…
705797active
Ladon
Ladon is a large-scale internal network penetration scanner written in C#, offering port scanning, service identification, network asset di…
295320active
imroc/req
Req is a simple yet powerful Go HTTP client library with chainable APIs, supporting HTTP/1.1, HTTP/2, and HTTP/3. It offers built-in debugg…
984853active
wasi-master/13ft
A self-hosted web service that bypasses paywalls on news sites by fetching pages as GoogleBot, a replacement for the defunct 12ft.io. It se…
834266active
hoothin/UserScripts
A collection of Greasemonkey/Tampermonkey userscripts by hoothin, including Pagetual (auto-pager infinite scrolling), Picviewer CE+ (online…
764264active
yacy/yacy_search_server
YaCy is a full search engine application in Java that combines a web crawler, a search index server, and a web front-end. It can run standa…
764018active
hardkoded/puppeteer-sharp
PuppeteerSharp is a .NET port of the official Node.js Puppeteer API for controlling headless or headful Chrome and Firefox. It supports nav…
993915active
s0md3v/XSStrike
XSStrike is a Python command-line Cross Site Scripting (XSS) detection suite that uses hand-written HTML/JavaScript parsers, context analys…
3115151maintenance
mxschmitt/playwright-go
Playwright for Go is a Go library that automates Chromium, Firefox, and WebKit browsers through a single API, supporting both headless and …
983481active
liyown/ai-trend-publish
TrendPublish is a TypeScript-based automated content pipeline for WeChat Official Accounts that scrapes multiple sources (Twitter/X, RSS, s…
823158active
Barabama/FreeNodes
A Python crawler that aggregates free proxy nodes (v2ray, Clash, vmess, vless, trojan, ss) from public websites and publishes them as auto-…
733096active
Virtual-Browser/VirtualBrowser
VirtualBrowser is a free, open-source anti-fingerprint browser built on Chromium that lets users create and manage multiple isolated browse…
933077active
aboul3la/Sublist3r
Sublist3r is a Python command-line tool that enumerates subdomains of a target domain using OSINT sources such as Google, Bing, Yahoo, Baid…
2311023maintenance
firecrawl/open-agent-builder
Open Agent Builder is a visual, no-code workflow builder for creating AI agent pipelines powered by Firecrawl web scraping. It provides a d…
382618active
hisxo/gitGraber
gitGraber is a Python3 command-line tool that monitors GitHub search results in real time to find leaked sensitive data such as API keys an…
662376active
probberechts/soccerdata
A Python library of scrapers that collect soccer data from popular websites like FBref, ESPN, WhoScored, Sofascore, SoFIFA, Understat, Club…
932040active
A9T9/RPA
Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible …
961985active
vmoranv/jshookmcp
An MCP server exposing 600+ tools across 34 domains for JavaScript reverse engineering and security research, including browser automation,…
591954active
tsingyuai/growth-lab
Growth Lab is an open-source, end-to-end AI growth system that runs product marketing and user-acquisition loops through natural-language c…
561861active
initstring/linkedin2username
A Python OSINT tool that scrapes LinkedIn employee lists for a target company and generates multiple probable username formats (e.g., first…
761825active
1N3/BlackWidow
BlackWidow is a Python-based web application spider that crawls a target site to collect URLs, dynamic parameters, subdomains, email addres…
571821active
bogdanfinn/tls-client
A Go HTTP client library with a net/http-like interface that lets you select specific browser TLS fingerprints (Chrome, Firefox, Safari, et…
901807active
Ge0rg3/requests-ip-rotator
A Python library that mounts AWS API Gateway as a proxy onto requests sessions, rotating source IPs on every request using AWS's large IP p…
901675active
collinbarrett/FilterLists
FilterLists is an independent, comprehensive web directory and REST API cataloging filter and host lists for blocking advertisements, track…
771645active
m3n0sd0n4ld/GooFuzz
GooFuzz is a Bash-based CLI tool that performs fuzzing-style reconnaissance using advanced Google searches (Google Dorking) via the Google …
571585active
AlisamTechnology/ATSCAN
ATSCAN is a Perl-based command-line scanner for mass dork searching and vulnerability exploitation. It combines search engine dorking with …
231583active
m8sec/CrossLinked
CrossLinked is a Python CLI tool that enumerates LinkedIn employee names for an organization by scraping search engine results, without nee…
231582active
hyperbrowserai/HyperAgent
HyperAgent is a TypeScript library and CLI that adds LLM-powered natural language commands to Playwright for browser automation. It support…
521540active
orangecoding/fredy
Fredy is a self-hosted Node.js application that continuously scrapes European real estate portals like ImmoScout24, Immowelt, Kleinanzeigen…
951444active
openwpm/OpenWPM
OpenWPM is a web privacy measurement framework built on Firefox with Selenium automation, designed to collect data from thousands to millio…
951417active
monosans/proxy-scraper-checker
A fast async Rust CLI tool that scrapes HTTP, SOCKS4, and SOCKS5 proxies from arbitrary text, HTML, or JSON sources, verifies each proxy by…
771318active
saeeddhqan/Maryam
OWASP Maryam is a modular open-source OSINT framework for harvesting data from open sources, search engines, and social networks. It provid…
101228active
hengliyin/cdfang-spider
A full-stack web application that scrapes Chengdu Housing Association lottery housing data and presents it through interactive charts and s…
531221active
ycdxsb/PocOrExp_in_Github
A Python CLI tool that automatically aggregates proof-of-concept (POC) and exploit (EXP) code from GitHub by CVE ID, using CVE information …
771198active
commons-app/apps-android-commons
The official community-maintained Wikimedia Commons Android app for uploading photos from an Android phone or tablet to Wikimedia Commons. …
951179active
austin-weeks/miasma
Miasma is a lightweight Rust web server that traps AI web scrapers in an endless pit of poisoned training data and self-referential links. …
821179active
WhiteNightShadow/hello_js_reverse_skill
An AI-powered 'Skill' package for JavaScript reverse engineering that plugs into AI coding tools like Claude Code, Cursor, and Codex. It pr…
781156active
sharebook-kr/pykrx
PyKrx is a Python library that scrapes stock and bond market data from the Korea Exchange (KRX) and Naver. It provides APIs for querying ti…
761084active
Cloxl/xhshow
A pure-algorithm Python library that generates Xiaohongshu (XHS/RedNote) request signature headers such as x-s, x-s-common, x-t, and x-rap-…
761052active
kelvinBen/AppInfoScanner
A Python-based static information-gathering scanner for mobile apps (Android APK/DEX, iOS IPA/Mach-O) and static web content (HTML, JS, H5)…
233554maintenance
apify/proxy-chain
A programmable HTTP/HTTPS proxy server library for Node.js, similar to Squid, with support for SSL/TLS, SOCKS4/5, authentication, upstream …
901020active
s-rah/onionscan
OnionScan is a free and open source Go CLI tool for investigating Tor hidden services (.onion sites) on the Dark Web. It scans sites for op…
233290maintenance
nickliqian/cnn_captcha
A Python project that uses convolutional neural networks built with TensorFlow to recognize character-based image captchas. It packages val…
322881maintenance
obheda12/GitDorker
GitDorker is a Python CLI tool that uses the GitHub Search API with a curated list of over 200 dorks to find sensitive information exposed …
322577maintenance
cyberagiinc/DevDocs
DevDocs is a free, private, UI-based MCP server that crawls and extracts technical documentation (using Crawl4AI and Playwright) and expose…
502106maintenance
670848654/SakuraAnime
A third-party Android client for the anime streaming sites Yhdm (Sakura Anime) and SiliSili, built in Java using jsoup for scraping site co…
102043maintenance
iam-abbas/Reddit-Stock-Trends
A Python application that scrapes Reddit via the PRAW API to identify trending stock tickers and analyzes their performance with yfinance. …
321590maintenance
BishopFox/GitGot
GitGot is a semi-automated, feedback-driven CLI tool for searching public GitHub data (code and gists) for exposed sensitive secrets. Users…
321571maintenance
gwen001/github-search
A collection of Python, PHP, and Bash scripts that perform targeted searches on GitHub via its search API to find secrets, keys, private re…
231511maintenance
ariya/phantomjs
PhantomJS is a headless WebKit browser scriptable with JavaScript, supporting page automation, screen capture, headless web testing, and ne…
1029449abandoned
clips/pattern
Pattern is a Python web mining module bundling tools for scraping (Google, Twitter, Wikipedia APIs, crawler, HTML DOM parser), NLP (POS tag…
668860abandoned
Nemo2011/bilibili-api
A Python SDK for programmatically accessing Bilibili (bilibili.com), including video data, danmaku comments, user info, and login. The repo…
104188abandoned
Greenwolf/social_mapper
Social Mapper is a Python 3 OSINT tool that enumerates and correlates social media profiles across sites like LinkedIn, Facebook, Twitter, …
324073abandoned
eth0izzle/shhgit
shhgit is a secrets detection tool that scans GitHub, GitLab, Bitbucket repositories and local directories for accidentally committed crede…
363977abandoned
the0demiurge/ShadowSocksShare
A Python/Flask web service that crawls shared Shadowsocks/ShadowsocksR accounts from public sharing websites, validates their connectivity,…
102992abandoned
Jinnrry/getAwayBSG
A Go CLI web crawler that scrapes job listings from Zhilian Zhaopin and housing (rental and second-hand) data from Lianjia across Chinese c…
101147abandoned

← prev page 6 / 6