domain: crawlers
581 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| scrapinghub/splash Splash is a lightweight, scriptable headless browser exposed as a service with an HTTP API, implemented in Python 3 using Twisted and Qt5. … | 32 | 4187 | maintenance |
| DataDog/guarddog GuardDog is a CLI tool from Datadog that identifies malicious packages on PyPI, npm, Go modules, Rust crates, RubyGems, GitHub Actions, and… | 99 | 1194 | active |
| constverum/ProxyBroker ProxyBroker is an asynchronous Python tool that finds public HTTP(S) and SOCKS4/5 proxies from ~50 sources and concurrently checks their ty… | 32 | 4159 | maintenance |
| xenova/chat-downloader Chat Downloader is a Python tool and library for retrieving chat messages from livestreams, videos, clips, and past broadcasts on platforms… | 45 | 1189 | active |
| intoli/user-agents A JavaScript/TypeScript npm package for generating random user agents weighted by real-world market share, with daily-updated data. It also… | 77 | 1188 | active |
| goodreasonai/ScrapeServ ScrapeServ is a self-hosted API service that accepts a URL and returns the website's data along with browser screenshots, using Playwright … | 25 | 1181 | active |
| rchipka/node-osmosis Osmosis is an HTML/XML parser and web scraper library for Node.js built on native libxml C bindings. It offers a chainable, promise-like in… | 32 | 4107 | maintenance |
| grangier/python-goose Python-Goose is a Python library that extracts the main body text, metadata, top image, and embedded videos from news article web pages. It… | 64 | 4106 | maintenance |
| online-judge-tools/oj A command-line tool that automates solving problems on online judges like AtCoder, Codeforces, and HackerRank. It downloads sample and syst… | 23 | 1169 | active |
| zu1k/proxypool A Go service that automatically crawls proxy nodes (ss, ssr, vmess, trojan) from Telegram channels, subscription URLs, and the public inter… | 23 | 4027 | maintenance |
| fanpei91/torsniff torsniff is a Go CLI tool that sniffs torrent metadata from the BitTorrent network by participating in the DHT and connecting to peers to d… | 10 | 4014 | maintenance |
| tholian-network/stealth Stealth is a secure, privacy-focused web browser, scraper, and proxy built in JavaScript that emphasizes automation, bandwidth efficiency, … | 32 | 1145 | active |
| datawhores/OF-Scraper OF-Scraper is a command-line tool for downloading media from OnlyFans and performing bulk actions like liking or unliking posts. It is a re… | 80 | 1143 | active |
| AndyTheFactory/newspaper4k Newspaper4k is a Python library and CLI for scraping and curating news articles, extracting text, titles, authors, publish dates, and metad… | 84 | 1140 | active |
| fwonggh/Bthub Bthub is a magnet link and torrent search engine, and this repository serves as its official address release page listing current and backu… | 76 | 1133 | active |
| webrecorder/browsertrix-crawler Browsertrix Crawler is a standalone browser-based high-fidelity web crawling system that runs in a single Docker container. It uses Puppete… | 99 | 1120 | active |
| platonai/Browser4 Browser4 is an AI-native browser engine built in Kotlin for autonomous agents, intelligent data extraction, and large-scale web automation.… | 100 | 1114 | active |
| elixir-crawly/crawly Crawly is a high-level web crawling and scraping framework for Elixir, modeled after Scrapy, where developers define spiders that fetch pag… | 37 | 1114 | active |
| nottelabs/reverse-api-engineer Reverse API Engineer is a Python CLI tool that captures browser network traffic (HAR) from a website and uses a configured AI model to gene… | 84 | 1113 | active |
| bellingcat/auto-archiver A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supp… | 92 | 1109 | active |
| vifreefly/kimuraframework Kimuraframework (Kimurai) is a Ruby web scraping framework with an AI-assisted DSL: an LLM generates XPath selectors from a schema on first… | 61 | 1102 | active |
| GPTaku Plugins insane-search is a Claude Code plugin that reads public web pages that would otherwise be blocked (403, CAPTCHA, WAF), escalating through p… | 60 | 1102 | active |
| Tyrrrz/YoutubeExplode A .NET library providing an abstraction layer over YouTube's internal API to query metadata for videos, playlists, and channels, and to res… | 99 | 3717 | maintenance |
| ruipgil/scraperjs Scraperjs is a Node.js web scraping library offering two scrapers: a lightweight StaticScraper using cheerio for static HTML, and a Dynamic… | 32 | 3714 | maintenance |
| techtanic/Discounted-Udemy-Course-Enroller A Python application (with GUI and CLI variants) that scrapes websites for 100% off Udemy course coupons and automatically enrolls the user… | 71 | 1081 | active |
| jmcarp/robobrowser RoboBrowser is a Pythonic library for browsing the web without a standalone browser, combining Requests for HTTP sessions with BeautifulSou… | 32 | 3692 | maintenance |
| jae-jae/fetcher-mcp A Model Context Protocol (MCP) server that fetches web page content using a Playwright headless browser, executing JavaScript to handle dyn… | 48 | 1076 | active |
| 1061700625/WeChat_Article A PyQt5 desktop application that crawls and downloads all articles from a specified WeChat official account. It uses Selenium to log in and… | 62 | 1075 | active |
| scrapfly/scrapfly-scrapers A collection of educational Python web scraping scripts for over 40 popular domains such as Amazon, AliExpress, BestBuy, and Twitter, built… | 74 | 1074 | active |
| soxoj/socid-extractor socid_extractor is a Python library and CLI that extracts structured account metadata and stable internal identifiers (usernames, UIDs, GAI… | 92 | 1073 | active |
| tomnomnom/assetfinder A Go command-line tool that discovers domains and subdomains potentially related to a given domain by querying multiple passive sources lik… | 23 | 3666 | maintenance |
| unitedstates/congress A community-run Python toolkit that collects and converts official U.S. Congress data—bills, amendments, roll call votes, nominations, and … | 52 | 1060 | active |
| cantino/selectorgadget SelectorGadget is an open-source bookmarklet and Chrome extension that generates CSS selectors for page elements through point-and-click se… | 69 | 1059 | active |
| spider-ios/autox-release AutoX is a desktop social media operations tool that automates one-click publishing of videos to multiple platforms such as Douyin, TikTok,… | 65 | 1059 | active |
| Henryhaohao/Bilibili_video_download A Python tool for downloading videos from Bilibili, supporting single-part and multi-part (分P) videos, bangumi episodes, and multiple downl… | 32 | 3583 | maintenance |
| jaimeiniesta/metainspector MetaInspector is a Ruby gem for web scraping that fetches a given URL and exposes its title, meta description, keywords, links, images, cha… | 69 | 1049 | active |
| caolvchong-top/twitter_download A Python command-line tool that scrapes and downloads images, videos (including GIFs), and text from Twitter/X user timelines. It supports … | 75 | 1045 | active |
| mandatoryprogrammer/thermoptic Thermoptic is a stealth HTTP proxy that routes requests from any HTTP client (like curl) through a containerized Chrome instance so the tra… | 51 | 1045 | active |
| turicas/brasil.io The backend of Brasil.IO, a platform that collects, cleans, and publishes Brazilian public open datasets in accessible formats. It automate… | 77 | 1044 | active |
| tmwgsicp/wechat-download-api An open-source API service for fetching WeChat official account articles, generating standard RSS 2.0 feeds, and exporting entire account a… | 77 | 1043 | active |
| lm-rebooter/NuggetsBooklet A Node.js script that downloads Juejin (掘金) booklet content for personal study by reusing the user's authenticated browser cookies. It expo… | 66 | 1041 | active |
| Anorov/cloudflare-scrape A Python module (cfscrape) built on Requests that bypasses Cloudflare's JavaScript anti-bot challenge page ('I'm Under Attack Mode') so scr… | 23 | 3538 | maintenance |
| jfilter/clean-text A Python package for cleaning and normalizing messy text, especially user-generated content from the web and social media. It fixes unicode… | 81 | 1027 | active |
| carcabot/tiktok-signature A self-hosted Node.js service that generates valid X-Bogus and X-Gnarly signature tokens for TikTok API requests using a headless browser r… | 93 | 1025 | active |
| wnma3mz/wechat_articles_spider A Python library for scraping WeChat Official Account articles, including article URLs, reading counts, likes, and comments. It can also do… | 23 | 3481 | maintenance |
| daijro/hrequests hrequests is a Python HTTP client library that replaces the requests library with browser TLS fingerprint replication, HTTP/2 support, and … | 30 | 1023 | active |
| JosephLai241/URS URS (Universal Reddit Scraper) is a comprehensive command-line tool written in Python (with Rust components) for scraping and archiving Red… | 64 | 1020 | active |
| owner888/phpspider phpspider is a PHP web crawling framework that lets developers build scrapers with a simple config array, handling multi-process workers, l… | 23 | 3462 | maintenance |
| wujunwei928/parse-video A Go library and CLI tool that parses short-video share links from 25+ Chinese platforms (Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo, e… | 81 | 1018 | active |
| Vinyzu/Botright Botright is a Python browser automation framework built on Playwright that provides undetectable, fingerprint-changing stealth browsing. It… | 76 | 1017 | active |
| pea3nut/Pxer Pxer is a userscript (installed via Tampermonkey) that acts as a crawler for pixiv.net, letting users batch-fetch artworks, collections, an… | 27 | 1010 | active |
| wreq wreq is an ergonomic, privacy-aware HTTP client written in Rust with a Python binding (wreq-python) that provides high-fidelity browser TLS… | 90 | 1006 | active |
| ranahaani/GNews GNews is a lightweight Python package that queries the Google News RSS feed and returns article results as usable JSON. It supports keyword… | 68 | 1006 | active |
| JoMingyu/google-play-scraper A Python library that provides APIs to crawl the Google Play Store for app details, reviews, and other data without any external dependenci… | 32 | 1006 | active |
| ma6254/FictionDown FictionDown is a Go-based command-line tool for batch downloading and crawling web novels from sites like Qidian and Biquge. It supports mu… | 24 | 1006 | active |
| jackwener/wechat-article-to-markdown A Python CLI tool that fetches WeChat Official Account articles using anti-detection browser automation (Camoufox) and converts them to cle… | 48 | 1002 | active |
| kevinzg/facebook-scraper A Python library for scraping public Facebook pages, groups, profiles, and posts without requiring an API key. It provides a simple get_pos… | 32 | 3270 | maintenance |
| wenbochang888/house A Java web scraper built with SpringBoot, HttpClient, and JSoup that crawls famous Tianya forum threads about China's housing market and co… | 32 | 3231 | maintenance |
| scrapy-plugins/scrapy-splash A Scrapy plugin that integrates the Splash headless browser service to enable crawling and scraping of JavaScript-rendered web pages. It pr… | 26 | 3227 | maintenance |
| ferventdesert/Hawk Hawk is a visual crawler and ETL IDE written in C#/WPF that lets users graphically scrape webpages, clean, transform, and store data withou… | 23 | 3213 | maintenance |
| CrawlScript/WebCollector WebCollector is an open-source Java web crawler framework that provides simple interfaces for building multi-threaded web crawlers quickly.… | 62 | 3083 | maintenance |
| jaeles-project/gospider GoSpider is a fast web spider/crawler written in Go that crawls sites in parallel and extracts URLs from sitemaps, robots.txt, JavaScript f… | 23 | 2993 | maintenance |
| kotartemiy/newscatcher A Python package that programmatically collects normalized news articles from thousands of news websites, filterable by topic, country, and… | 32 | 2987 | maintenance |
| Threezh1/JSFinder JSFinder is a Python command-line tool that crawls a website's JavaScript files and extracts URLs and subdomains using regex parsing. It su… | 32 | 2976 | maintenance |
| facundoolano/google-play-scraper A Node.js library that scrapes application data from the Google Play store, exposing methods for app details, search, reviews, permissions,… | 66 | 2949 | maintenance |
| howie6879/owllook owllook is a self-hosted vertical search engine for Chinese web novels, built on Python with Sanic, MongoDB, and Redis. It aggregates resul… | 32 | 2871 | maintenance |
| CharlesPikachu/DecryptLogin A Python library providing programmatic login APIs for popular websites (Weibo, Bilibili, Zhihu, GitHub, Taobao, etc.) built on the request… | 32 | 2855 | maintenance |
| yann-shi/dht A Go library implementing the BitTorrent DHT protocol (BEP-3, 5, 9, 10) with two modes: a standard DHT server and a crawling mode for harve… | 23 | 2769 | maintenance |
| brianway/webporter webporter is a Java crawler application built on the webmagic framework that demonstrates a complete pipeline of data crawling, persistence… | 23 | 2766 | maintenance |
| DormyMo/SpiderKeeper SpiderKeeper is a self-hosted web-based admin dashboard for managing Scrapy spiders running on Scrapyd servers. It provides spider scheduli… | 32 | 2763 | maintenance |
| jeanphix/Ghost.py Ghost.py is a scriptable WebKit-based web client library for Python, built on PySide2/Qt5, allowing programmatic page loading and content i… | 32 | 2756 | maintenance |
| luin/readability A Node.js library that extracts clean, readable article content from any web page, based on arc90's readability project. It returns the art… | 32 | 2519 | maintenance |
| xtuhcy/gecco Gecco is a lightweight, easy-to-use web crawler framework for Java that lets developers define crawlers with annotation-based jQuery-style … | 51 | 2510 | maintenance |
| lorien/grab Grab is a Python web scraping framework providing HTTP request handling, proxy/cookie support, and XPath-based HTML parsing, plus a Spider … | 51 | 2463 | maintenance |
| decaywood/XueQiuSuperSpider A Java 8 web scraping framework for collecting stock data from Xueqiu (Snowball) and other Chinese financial sites. It is built around comp… | 32 | 2433 | maintenance |
| paquettg/php-html-parser A PHP library that parses HTML into a DOM and lets you find and manipulate tags using CSS selectors, similar to jQuery. It is designed for … | 32 | 2398 | maintenance |
| QianyanTech/Image-Downloader A Python application that crawls and downloads images from Google, Bing, and Baidu using Selenium or API drivers. It offers both a PyQt5 GU… | 23 | 2362 | maintenance |
| lucasjinreal/weibo_terminater A Python-based web scraper that crawls Weibo (Sina's microblog platform) to collect user posts, comments, followers, and conversation pairs… | 32 | 2317 | maintenance |
| ageitgey/node-unfluff A Node.js library and CLI tool that automatically extracts the main body content and metadata (title, author, date, images, tags, links) fr… | 32 | 2158 | maintenance |
| t9tio/cloudquery CloudQuery is a tool that turns any website into a JSON API by fetching pages with headless Chrome and extracting data via CSS selectors. I… | 32 | 2149 | maintenance |
| php-embed/Embed A PHP library that extracts metadata and embed information from any web page or web service using oEmbed, OpenGraph, Twitter Cards, and HTM… | 93 | 2140 | maintenance |
| anouarbensaad/vulnx VulnX is a Python CLI tool that detects CMS types (WordPress, Joomla, Drupal, etc.), gathers target information like subdomains and DNS rec… | 23 | 2138 | maintenance |
| PuerkitoBio/gocrawl gocrawl is a polite, slim and concurrent web crawler library written in Go. It respects robots.txt rules, applies per-host crawl delays, an… | 23 | 2052 | maintenance |
| Nekmo/dirhunt Dirhunt is a Python CLI web crawler optimized for finding and analyzing web directories without brute-forcing paths. It detects 'index of' … | 23 | 2007 | maintenance |
| awolfly9/IPProxyTool A Python/Scrapy application that crawls free proxy websites, validates the collected proxy IPs against target sites, and stores usable prox… | 32 | 1998 | maintenance |
| Xyntax/POC-T POC-T is a Python 2.7 plugin-based concurrent framework for penetration testing tasks such as crawling, bruteforcing, and batch PoC/EXP ver… | 23 | 1936 | maintenance |
| scrapy/scrapely Scrapely is a pure-Python library for extracting structured data from HTML pages. It learns a parser from example pages annotated with the … | 32 | 1883 | maintenance |
| xianhu/PSpider PSpider is a simple, easy-to-read web spider framework written in Python 3.8+. It uses a Fetcher/Parser/Saver pipeline with queues for mult… | 32 | 1835 | maintenance |
| reworkd/tarsier Tarsier is a Python library providing vision utilities for LLM-driven web interaction agents. It visually tags interactable page elements w… | 17 | 1761 | maintenance |
| howie6879/ruia Ruia is an async web scraping micro-framework for Python 3.6+ built on asyncio and aiohttp. It offers declarative Item/Field extraction (XP… | 23 | 1738 | maintenance |
| henson/proxypool A Golang IP proxy pool service that scrapes free proxy sources, validates them, stores them in a database, and exposes a JSON API for crawl… | 32 | 1701 | maintenance |
| YoongiKim/AutoCrawler A Python multiprocess image web crawler that downloads images from Google and Naver image search using Selenium and ChromeDriver. It is des… | 32 | 1691 | maintenance |
| th3unkn0n/TeleGram-Scraper A Python command-line tool that scrapes Telegram groups and exports all member information to CSV using the Telegram API. It also includes … | 10 | 1672 | maintenance |
| aivarsk/scrapy-proxies A Scrapy downloader middleware that routes requests through random proxies from a configurable list to avoid IP bans. It supports multiple … | 32 | 1667 | maintenance |
| gigablast/open-source-search-engine Gigablast is a distributed open source web and enterprise search engine with a built-in spider/crawler, written in C/C++ for Linux. It powe… | 32 | 1601 | maintenance |
| propublica/upton Upton is a Ruby framework that handles the repetitive parts of writing web scrapers, letting developers focus on site-specific CSS selector… | 32 | 1597 | maintenance |
| 0xHJK/dumpall dumpall is a Python command-line tool for exploiting information disclosure vulnerabilities on web servers. It reconstructs source code fro… | 23 | 1579 | maintenance |
| lqqyt2423/wechat_spider A Node.js WeChat crawler that uses a man-in-the-middle proxy (AnyProxy) to batch-collect WeChat official account article data, including co… | 76 | 1574 | maintenance |
| github/lightcrawler Lightcrawler is a Node.js CLI tool that crawls a website by following links and runs each discovered page through Google Lighthouse audits.… | 10 | 1564 | maintenance |
| headzoo/surf Surf is a Go library that implements a stateful virtual web browser controlled programmatically. It supports cookies, history, bookmarks, u… | 23 | 1544 | maintenance |