domain: crawlers
581 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| MiningCattiva/x-spider X-Spider is a desktop media downloader for X (Twitter) that fetches images and videos from accounts. It supports media filters, configurabl… | 10 | 1261 | abandoned |
| Zeal-L/BiliBili-Manga-Downloader A GUI-based Bilibili Manga downloader built with Python and PySide6 that supports keyword search, QR-code login, multi-threaded downloads, … | 36 | 1242 | abandoned |
| syrusakbary/gdom GDOM is a Python CLI tool that lets you scrape and traverse web page DOM using GraphQL queries, built on the Graphene framework. You write … | 32 | 1242 | abandoned |
| chenjiandongx/mzitu A Python web crawler that downloads full photo galleries from mzitu.com, collecting over 3,000 image sets. It also performs word-frequency … | 32 | 1220 | abandoned |
| k1995/BaiduyunSpider A distributed crawler and search engine for Baidu Netdisk (Baidu Cloud) shared links, built with Scrapy and Scrapy-Redis, with a React-base… | 23 | 1177 | abandoned |
| xiaoguyu/wechatDownload A desktop application built with Electron, TypeScript, and Vue3 for downloading WeChat Official Account (公众号) articles. It captures require… | 10 | 1170 | abandoned |
| shadowmoose/RedditDownloader Reddit Media Downloader is a Python CLI tool that scans Reddit posts, comments, saved lists, and subreddits to download linked media locall… | 10 | 1161 | abandoned |
| xillwillx/skiptracer Skiptracer is a Python-based OSINT web scraping framework that aggregates paid and free lookup services to enumerate target information suc… | 23 | 1146 | abandoned |
| ring04h/weakfilescan A Python-based multi-threaded sensitive information leakage detection tool that crawls a target site, dynamically builds dictionary rules f… | 32 | 1138 | abandoned |
| joshhighet/ransomwatch A transparent ransomware claim tracker that monitors ransomware groups' leak sites on the dark web and records their victim claims. It aggr… | 10 | 1116 | abandoned |
| bytebuff/JSpider JSpider is a collection of JavaScript decryption files for website-encrypted request parameters, shared weekly alongside Python snippets th… | 32 | 1091 | abandoned |
| VikParuchuri/apartment-finder A Python Slack bot that scrapes Craigslist for real-time apartment listings matching configurable criteria (price, neighborhoods, transit p… | 10 | 1059 | abandoned |
| 7sDream/zhihu-py3 An unofficial Python 3 API library for Zhihu, the Chinese Q&A site, letting users build objects from Zhihu URLs to fetch questions, answers… | 10 | 1033 | abandoned |
| SpiderClub/smart_login A Python collection of simulated login implementations for major Chinese websites (Weibo, Zhihu, QQ Zone, JD, Baidu, etc.), using either di… | 32 | 1008 | abandoned |
| puppeteer/puppeteer Puppeteer is a JavaScript library providing a high-level API to control Chrome or Firefox over the DevTools Protocol or WebDriver BiDi, run… | 95 | 95497 | stable |
| lightpanda-io/browser Lightpanda is a headless web browser built from scratch in Zig, designed for AI agents, automation, and web scraping rather than human rend… | 97 | 34269 | active |
| cheeriojs/cheerio Cheerio is a fast, flexible JavaScript library for parsing and manipulating HTML and XML with a jQuery-like API. It works in both browser a… | 84 | 30468 | stable |
| Skyvern-AI/skyvern Skyvern is an open-source AI browser automation framework that uses LLMs and computer vision to interact with websites, offering a Playwrig… | 88 | 22852 | active |
| AutomaApp/automa Automa is a browser extension for Chrome and Firefox that lets users automate browser tasks by visually connecting blocks into workflows. I… | 66 | 21586 | active |
| chromedp chromedp is a Go library for driving Chrome and other browsers via the Chrome DevTools Protocol, with no external dependencies. It supports… | 84 | 13264 | active |
| seleniumbase/SeleniumBase SeleniumBase is an all-in-one Python browser automation framework for web testing, crawling, and scraping, built on Selenium/WebDriver with… | 95 | 12955 | active |
| TeamWiseFlow/xiaobei Xiaobei is an open-source multi-agent system that automates social media marketing and customer acquisition for solo entrepreneurs and smal… | 90 | 8463 | active |
| adithya-s-k/omniparse OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into… | 49 | 7815 | active |
| go-rod/rod Rod is a high-level Go library that drives Chrome via the Chrome DevTools Protocol for web automation and scraping. It offers both high-lev… | 66 | 7077 | active |
| lwthiker/curl-impersonate A special build of curl (and libcurl) whose TLS and HTTP/2 handshakes are identical to real browsers like Chrome, Edge, Safari, and Firefox… | 23 | 6895 | active |
| aidlearning/AidLearning-FrameWork AidLux (originally AidLearning) is an AIoT development platform that runs a native Ubuntu Linux environment with GUI, deep learning tooling… | 70 | 5797 | active |
| Ladon Ladon is a large-scale internal network penetration scanner written in C#, offering port scanning, service identification, network asset di… | 29 | 5320 | active |
| imroc/req Req is a simple yet powerful Go HTTP client library with chainable APIs, supporting HTTP/1.1, HTTP/2, and HTTP/3. It offers built-in debugg… | 98 | 4853 | active |
| wasi-master/13ft A self-hosted web service that bypasses paywalls on news sites by fetching pages as GoogleBot, a replacement for the defunct 12ft.io. It se… | 83 | 4266 | active |
| hoothin/UserScripts A collection of Greasemonkey/Tampermonkey userscripts by hoothin, including Pagetual (auto-pager infinite scrolling), Picviewer CE+ (online… | 76 | 4264 | active |
| yacy/yacy_search_server YaCy is a full search engine application in Java that combines a web crawler, a search index server, and a web front-end. It can run standa… | 76 | 4018 | active |
| hardkoded/puppeteer-sharp PuppeteerSharp is a .NET port of the official Node.js Puppeteer API for controlling headless or headful Chrome and Firefox. It supports nav… | 99 | 3915 | active |
| s0md3v/XSStrike XSStrike is a Python command-line Cross Site Scripting (XSS) detection suite that uses hand-written HTML/JavaScript parsers, context analys… | 31 | 15151 | maintenance |
| mxschmitt/playwright-go Playwright for Go is a Go library that automates Chromium, Firefox, and WebKit browsers through a single API, supporting both headless and … | 98 | 3481 | active |
| liyown/ai-trend-publish TrendPublish is a TypeScript-based automated content pipeline for WeChat Official Accounts that scrapes multiple sources (Twitter/X, RSS, s… | 82 | 3158 | active |
| Barabama/FreeNodes A Python crawler that aggregates free proxy nodes (v2ray, Clash, vmess, vless, trojan, ss) from public websites and publishes them as auto-… | 73 | 3096 | active |
| Virtual-Browser/VirtualBrowser VirtualBrowser is a free, open-source anti-fingerprint browser built on Chromium that lets users create and manage multiple isolated browse… | 93 | 3077 | active |
| aboul3la/Sublist3r Sublist3r is a Python command-line tool that enumerates subdomains of a target domain using OSINT sources such as Google, Bing, Yahoo, Baid… | 23 | 11023 | maintenance |
| firecrawl/open-agent-builder Open Agent Builder is a visual, no-code workflow builder for creating AI agent pipelines powered by Firecrawl web scraping. It provides a d… | 38 | 2618 | active |
| hisxo/gitGraber gitGraber is a Python3 command-line tool that monitors GitHub search results in real time to find leaked sensitive data such as API keys an… | 66 | 2376 | active |
| probberechts/soccerdata A Python library of scrapers that collect soccer data from popular websites like FBref, ESPN, WhoScored, Sofascore, SoFIFA, Understat, Club… | 93 | 2040 | active |
| A9T9/RPA Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible … | 96 | 1985 | active |
| vmoranv/jshookmcp An MCP server exposing 600+ tools across 34 domains for JavaScript reverse engineering and security research, including browser automation,… | 59 | 1954 | active |
| tsingyuai/growth-lab Growth Lab is an open-source, end-to-end AI growth system that runs product marketing and user-acquisition loops through natural-language c… | 56 | 1861 | active |
| initstring/linkedin2username A Python OSINT tool that scrapes LinkedIn employee lists for a target company and generates multiple probable username formats (e.g., first… | 76 | 1825 | active |
| 1N3/BlackWidow BlackWidow is a Python-based web application spider that crawls a target site to collect URLs, dynamic parameters, subdomains, email addres… | 57 | 1821 | active |
| bogdanfinn/tls-client A Go HTTP client library with a net/http-like interface that lets you select specific browser TLS fingerprints (Chrome, Firefox, Safari, et… | 90 | 1807 | active |
| Ge0rg3/requests-ip-rotator A Python library that mounts AWS API Gateway as a proxy onto requests sessions, rotating source IPs on every request using AWS's large IP p… | 90 | 1675 | active |
| collinbarrett/FilterLists FilterLists is an independent, comprehensive web directory and REST API cataloging filter and host lists for blocking advertisements, track… | 77 | 1645 | active |
| m3n0sd0n4ld/GooFuzz GooFuzz is a Bash-based CLI tool that performs fuzzing-style reconnaissance using advanced Google searches (Google Dorking) via the Google … | 57 | 1585 | active |
| AlisamTechnology/ATSCAN ATSCAN is a Perl-based command-line scanner for mass dork searching and vulnerability exploitation. It combines search engine dorking with … | 23 | 1583 | active |
| m8sec/CrossLinked CrossLinked is a Python CLI tool that enumerates LinkedIn employee names for an organization by scraping search engine results, without nee… | 23 | 1582 | active |
| hyperbrowserai/HyperAgent HyperAgent is a TypeScript library and CLI that adds LLM-powered natural language commands to Playwright for browser automation. It support… | 52 | 1540 | active |
| orangecoding/fredy Fredy is a self-hosted Node.js application that continuously scrapes European real estate portals like ImmoScout24, Immowelt, Kleinanzeigen… | 95 | 1444 | active |
| openwpm/OpenWPM OpenWPM is a web privacy measurement framework built on Firefox with Selenium automation, designed to collect data from thousands to millio… | 95 | 1417 | active |
| monosans/proxy-scraper-checker A fast async Rust CLI tool that scrapes HTTP, SOCKS4, and SOCKS5 proxies from arbitrary text, HTML, or JSON sources, verifies each proxy by… | 77 | 1318 | active |
| saeeddhqan/Maryam OWASP Maryam is a modular open-source OSINT framework for harvesting data from open sources, search engines, and social networks. It provid… | 10 | 1228 | active |
| hengliyin/cdfang-spider A full-stack web application that scrapes Chengdu Housing Association lottery housing data and presents it through interactive charts and s… | 53 | 1221 | active |
| ycdxsb/PocOrExp_in_Github A Python CLI tool that automatically aggregates proof-of-concept (POC) and exploit (EXP) code from GitHub by CVE ID, using CVE information … | 77 | 1198 | active |
| commons-app/apps-android-commons The official community-maintained Wikimedia Commons Android app for uploading photos from an Android phone or tablet to Wikimedia Commons. … | 95 | 1179 | active |
| austin-weeks/miasma Miasma is a lightweight Rust web server that traps AI web scrapers in an endless pit of poisoned training data and self-referential links. … | 82 | 1179 | active |
| WhiteNightShadow/hello_js_reverse_skill An AI-powered 'Skill' package for JavaScript reverse engineering that plugs into AI coding tools like Claude Code, Cursor, and Codex. It pr… | 78 | 1156 | active |
| sharebook-kr/pykrx PyKrx is a Python library that scrapes stock and bond market data from the Korea Exchange (KRX) and Naver. It provides APIs for querying ti… | 76 | 1084 | active |
| Cloxl/xhshow A pure-algorithm Python library that generates Xiaohongshu (XHS/RedNote) request signature headers such as x-s, x-s-common, x-t, and x-rap-… | 76 | 1052 | active |
| kelvinBen/AppInfoScanner A Python-based static information-gathering scanner for mobile apps (Android APK/DEX, iOS IPA/Mach-O) and static web content (HTML, JS, H5)… | 23 | 3554 | maintenance |
| apify/proxy-chain A programmable HTTP/HTTPS proxy server library for Node.js, similar to Squid, with support for SSL/TLS, SOCKS4/5, authentication, upstream … | 90 | 1020 | active |
| s-rah/onionscan OnionScan is a free and open source Go CLI tool for investigating Tor hidden services (.onion sites) on the Dark Web. It scans sites for op… | 23 | 3290 | maintenance |
| nickliqian/cnn_captcha A Python project that uses convolutional neural networks built with TensorFlow to recognize character-based image captchas. It packages val… | 32 | 2881 | maintenance |
| obheda12/GitDorker GitDorker is a Python CLI tool that uses the GitHub Search API with a curated list of over 200 dorks to find sensitive information exposed … | 32 | 2577 | maintenance |
| cyberagiinc/DevDocs DevDocs is a free, private, UI-based MCP server that crawls and extracts technical documentation (using Crawl4AI and Playwright) and expose… | 50 | 2106 | maintenance |
| 670848654/SakuraAnime A third-party Android client for the anime streaming sites Yhdm (Sakura Anime) and SiliSili, built in Java using jsoup for scraping site co… | 10 | 2043 | maintenance |
| iam-abbas/Reddit-Stock-Trends A Python application that scrapes Reddit via the PRAW API to identify trending stock tickers and analyzes their performance with yfinance. … | 32 | 1590 | maintenance |
| BishopFox/GitGot GitGot is a semi-automated, feedback-driven CLI tool for searching public GitHub data (code and gists) for exposed sensitive secrets. Users… | 32 | 1571 | maintenance |
| gwen001/github-search A collection of Python, PHP, and Bash scripts that perform targeted searches on GitHub via its search API to find secrets, keys, private re… | 23 | 1511 | maintenance |
| ariya/phantomjs PhantomJS is a headless WebKit browser scriptable with JavaScript, supporting page automation, screen capture, headless web testing, and ne… | 10 | 29449 | abandoned |
| clips/pattern Pattern is a Python web mining module bundling tools for scraping (Google, Twitter, Wikipedia APIs, crawler, HTML DOM parser), NLP (POS tag… | 66 | 8860 | abandoned |
| Nemo2011/bilibili-api A Python SDK for programmatically accessing Bilibili (bilibili.com), including video data, danmaku comments, user info, and login. The repo… | 10 | 4188 | abandoned |
| Greenwolf/social_mapper Social Mapper is a Python 3 OSINT tool that enumerates and correlates social media profiles across sites like LinkedIn, Facebook, Twitter, … | 32 | 4073 | abandoned |
| eth0izzle/shhgit shhgit is a secrets detection tool that scans GitHub, GitLab, Bitbucket repositories and local directories for accidentally committed crede… | 36 | 3977 | abandoned |
| the0demiurge/ShadowSocksShare A Python/Flask web service that crawls shared Shadowsocks/ShadowsocksR accounts from public sharing websites, validates their connectivity,… | 10 | 2992 | abandoned |
| Jinnrry/getAwayBSG A Go CLI web crawler that scrapes job listings from Zhilian Zhaopin and housing (rental and second-hand) data from Lianjia across Chinese c… | 10 | 1147 | abandoned |
← prev page 6 / 6