domain: crawlers
581 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| chaitin/rad Rad (Radium) is a browser-based web crawler built for security scanning, driving a real Chrome browser to discover URLs and requests across… | 23 | 1514 | maintenance |
| VideoData/DY-Data A collection of Douyin (Chinese TikTok) scraping tools and API source code covering search, user, video, live stream, comments, danmaku, an… | 32 | 1510 | maintenance |
| fossasia/event-collect A Python CLI tool that scrapes event website listings (e.g., EventBrite search results) and converts them into the Open Event JSON format. … | 10 | 1505 | maintenance |
| 78/ssbc The source code of the Shousibaocai (手撕包菜) website, a self-hosted BitTorrent magnet link search engine. It combines a Node.js DHT network s… | 32 | 1476 | maintenance |
| AlexCSDev/PatreonDownloader A command-line application for downloading content posted by creators on patreon.com, including posts, attachments, descriptions, and embed… | 69 | 1471 | maintenance |
| udacimak/udacimak Udacimak is a command-line tool that downloads Udacity Nanodegree and course content, including videos and materials, and renders them loca… | 32 | 1457 | maintenance |
| liguobao/HouseSearch A map-driven rental housing aggregation platform that continuously crawls public rental listings from sources like Douban, Beike, and Xiaoh… | 76 | 1439 | maintenance |
| xisuo67/XHS-Spider XHS-Spider is a polished Windows desktop (WPF, .NET 6) tool for collecting Xiaohongshu (Little Red Book) data, including keyword and user s… | 64 | 1435 | maintenance |
| damklis/DataEngineeringProject An end-to-end data engineering project that scrapes news from RSS feeds via Airflow-scheduled Python scrapers and streams them through Kafk… | 32 | 1429 | maintenance |
| jonnnnyw/php-phantomjs A PHP library that wraps the PhantomJS headless browser, letting PHP applications load web pages with full JavaScript support and inspect t… | 23 | 1426 | maintenance |
| kotartemiy/pygooglenews A Python wrapper around the Google News RSS feed providing top stories, topic and geolocation feeds, and full-text search with date-range a… | 32 | 1391 | maintenance |
| PKUJohnson/OpenData OpenDataTools is a Python library that scrapes financial and investment data from various websites and exposes it through simple, easy-to-u… | 32 | 1367 | maintenance |
| martinsbalodis/web-scraper-chrome-extension Web Scraper is a Chrome browser extension for extracting data from web pages without writing code. Users define sitemaps describing how to … | 32 | 1363 | maintenance |
| felipecsl/wombat Wombat is a lightweight Ruby web crawler and scraper library with an elegant DSL for extracting structured data from web pages. It lets dev… | 73 | 1360 | maintenance |
| ai-to-ai/Auto-Gmail-Creator A Python Selenium-based bot that bulk-creates Gmail accounts automatically, using sms-activate.org for phone verification and webdriver-man… | 31 | 1360 | maintenance |
| huaying/instagram-crawler A Python command-line crawler that scrapes Instagram posts, profiles, and hashtag data using Selenium and ChromeDriver, without the officia… | 32 | 1350 | maintenance |
| adriancooney/puppeteer-heap-snapshot A TypeScript library and CLI for capturing Chrome DevTools heap snapshots from Puppeteer-controlled pages and querying them for objects wit… | 23 | 1350 | maintenance |
| scrapinghub/frontera Frontera is a Python web crawling framework that implements a scalable crawl frontier, storing and prioritizing links extracted by crawlers… | 34 | 1332 | maintenance |
| laramies/metagoofil Metagoofil is a Python command-line OSINT tool that searches Google for public documents (pdf, doc, xls, ppt) on target websites, downloads… | 32 | 1313 | maintenance |
| c4tcom/Katana Katana-ds is a Python CLI tool that automates advanced Google queries known as Google Dorks (Google Dorking), with optional Tor support for… | 10 | 1304 | maintenance |
| InstaPy/instagram-profilecrawl A Python script from the InstaPy ecosystem that crawls public Instagram profile information such as post counts, follower counts, and post … | 32 | 1301 | maintenance |
| devanshbatham/FavFreak A Python CLI tool that fetches favicon.ico files from lists of URLs, computes their mmh3 hashes, and groups domains/subdomains/IPs by match… | 23 | 1298 | maintenance |
| sunra/php-simple-html-dom-parser A PHP library adapting Simple HTML DOM Parser for Composer and PSR-0, allowing easy HTML parsing and manipulation. It supports invalid HTML… | 23 | 1281 | maintenance |
| Sniper970119/dianping_spider A Python web scraper for Dianping (大众点评) that extracts search results, shop details, and reviews, handling dynamic font encryption without … | 72 | 1278 | maintenance |
| dragnet-org/dragnet Dragnet is a Python library that uses machine learning models to extract the main article content, and optionally user comments, from HTML … | 36 | 1274 | maintenance |
| useragents/Zefoy-TikTok-Automator A Python/Selenium automation script that drives zefoy.com to send TikTok followers, views, likes, shares, favorites, and comment likes. It … | 32 | 1265 | maintenance |
| kiddyuchina/Beanbun Beanbun is a multi-process web crawler framework written in PHP, built on Workerman with Guzzle as the default downloader. It supports dist… | 61 | 1259 | maintenance |
| dchrastil/ScrapedIn A Python CLI tool that scrapes LinkedIn without API restrictions to enumerate employees of a target company for red team or social engineer… | 32 | 1235 | maintenance |
| istresearch/scrapy-cluster Scrapy Cluster is a distributed web scraping framework built on Scrapy that uses Redis to coordinate crawl requests and Kafka as a data bus… | 10 | 1225 | maintenance |
| vysecurity/LinkedInt LinkedInt is a Python CLI tool for LinkedIn reconnaissance that scrapes employee profiles for a target company and generates an HTML report… | 10 | 1214 | maintenance |
| Mapaler/PixivUserBatchDownload A userscript (browser extension-style script) that batch-downloads all public works of a Pixiv artist directly from the Pixiv website, send… | 47 | 1191 | maintenance |
| timwhitez/crawlergo_x_XRAY A Python glue script that combines the crawlergo dynamic crawler with the XRAY passive vulnerability scanner, replaying crawled URLs throug… | 32 | 1182 | maintenance |
| yutto-dev/bilili bilili is a Python CLI tool for downloading Bilibili videos (including bangumi/season content) along with danmaku comments and subtitles. I… | 10 | 1180 | maintenance |
| WebSpiderUtils/verification_code A research repository documenting approaches and code for solving mainstream CAPTCHA systems such as Geetest, NetEase Yidun, and Aliyun CAP… | 32 | 1165 | maintenance |
| dixudx/tumblr-crawler A Python script that downloads all photos and videos from specified Tumblr blogs. It supports batch site lists, proxy configuration, and sk… | 59 | 1157 | maintenance |
| holgerd77/django-dynamic-scraper A Django app that builds on the Scrapy framework and lets you create and manage Scrapy spiders through the Django admin interface. It remov… | 23 | 1157 | maintenance |
| s045pd/DarkNet_ChineseTrading A Python-based real-time crawler that monitors Chinese-language darknet marketplaces over Tor, with automatic account registration, login, … | 10 | 1149 | maintenance |
| Jinnrry/RobotHelper RobotHelper is an Android automation script framework written in Java, providing common building blocks like screen capture, image-based po… | 23 | 1136 | maintenance |
| juancarlospaco/faster-than-requests A Python 3 HTTP client library claiming to be much faster than the popular Requests library, implemented with a compiled core and minimal d… | 67 | 1129 | maintenance |
| bonfy/github-trending A Python script that scrapes GitHub's trending page daily and archives the most popular repositories as Markdown files. It is designed to r… | 77 | 1128 | maintenance |
| kohlschutter/boilerpipe boilerpipe is a Java library for removing boilerplate (ads, navigation, headers) from HTML pages and extracting the main full text content.… | 32 | 1127 | maintenance |
| medialab/artoo artoo.js is a JavaScript library injected into a webpage's context (typically via a bookmarklet) that provides client-side web scraping uti… | 23 | 1119 | maintenance |
| exorde-labs/exorde-client The Exorde client is a Python CLI worker node for the Exorde Network, a decentralized protocol where participants scrape social media and w… | 60 | 1085 | maintenance |
| acikyazilimagi/afet-org An open-source earthquake relief platform (depremyardim.com / afetharita.com) that aggregates calls for help from Twitter, WhatsApp, Telegr… | 31 | 1079 | maintenance |
| jonbakerfish/TweetScraper TweetScraper is a Scrapy-based crawler that scrapes tweets and user information from Twitter Search without using Twitter's official APIs. … | 23 | 1062 | maintenance |
| utkarshkukreti/select.rs A Rust library for extracting useful data from HTML documents, built around predicate-based DOM queries. It is designed for web scraping ta… | 38 | 1020 | maintenance |
| Algebra-FUN/WeReadScan A Python library that uses Selenium headless browsers to scan purchased books from WeRead (WeChat Reading) and convert them into local PDF … | 32 | 1002 | maintenance |
| JSREI/ast-hook-for-js-RE A browser memory roaming tool for JavaScript reverse engineering that hooks variable assignments via AST-transformed proxy responses. It le… | 23 | 1912 | experimental |
| tinyfish-io/bigset-oss BigSet is a self-hostable application that turns a natural-language sentence into a structured, regularly refreshed dataset by dispatching … | 76 | 1684 | experimental |
| kkyon/botflow Botflow is a Python dataflow programming framework for building data pipelines using pipes and routes, with parallelism via coroutines and … | 62 | 1196 | experimental |
| oxylabs/ai-scraper-py AI-Scraper is a Python library and scrape agent from Oxylabs AI Studio that extracts data from webpages using natural language prompts inst… | 51 | 1055 | experimental |
| twintproject/twint Twint is a Python CLI tool and library that scrapes tweets, followers, following, and likes from Twitter without using the official API or … | 10 | 16398 | abandoned |
| wechat-article/wechat-article-exporter An online batch downloader for WeChat Official Account articles that exports posts in HTML, JSON, Excel, TXT, Markdown, and DOCX formats, i… | 68 | 12789 | abandoned |
| xiandanin/magnetW MagnetW is a cross-platform desktop application built with Electron and Vue that aggregates magnet link search results from multiple torren… | 10 | 11271 | abandoned |
| yangyangwithgnu/hardseed hardseed is a C++ command-line tool that scrapes adult forum threads (aicheng, caoliu) to batch-download images and torrent seed files, wit… | 32 | 9198 | abandoned |
| TonyChen56/WeChatRobot A C++ WeChat hooking toolkit and robot framework providing wxhook APIs, WeChat database decryption, and Official Account (公众号) scraping, pl… | 58 | 7188 | abandoned |
| xchaoinfo/fuck-login A Python library that simulates programmatic login to popular Chinese websites like Zhihu, Weibo, Baidu, and Douban, built on requests, Pil… | 10 | 5870 | abandoned |
| hanc00l/wooyun_public A crawler and search application for the archived Wooyun.org security vulnerability disclosure platform, containing ~40k-88k public vulnera… | 10 | 4399 | abandoned |
| qiyeboy/IPProxyPool IPProxyPool is a Python proxy pool service that crawls free proxy IPs from the web, validates them, stores them in a database (SQLite by de… | 23 | 4286 | abandoned |
| bisguzar/twitter-scraper A Python library that scrapes Twitter's frontend JavaScript API without authentication, letting users fetch tweets from profiles or hashtag… | 10 | 4005 | abandoned |
| pyppeteer/pyppeteer Pyppeteer is an unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. It supports page navigati… | 32 | 3942 | abandoned |
| miyakogi/pyppeteer An unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. This original repository has moved to … | 10 | 3550 | abandoned |
| bowenpay/wechat-spider A Python-based web crawler for scraping articles from WeChat public accounts (微信公众号), built on Django with MySQL and Redis, including a web… | 32 | 3369 | abandoned |
| LiuXingMing/SinaSpider A Python web crawler for Sina Weibo (Chinese microblog) built on Scrapy, with three versions: a standalone spider, a distributed version us… | 32 | 3285 | abandoned |
| gnemoug/distribute_crawler A distributed web crawler built on Scrapy, Redis, MongoDB, and Graphite, demonstrated with a spider for a Chinese book-download site. Redis… | 32 | 3238 | abandoned |
| harismuneer/Ultimate-Social-Scrapers A collection of Python-based scraping tools that extract public data from Facebook, Instagram, and Twitter (X), including posts, media, fol… | 44 | 3151 | abandoned |
| airingursb/bilibili-user A Python web crawler that scrapes Bilibili user profiles (id, nickname, gender, avatar, level, birthday, location, etc.) and stores them in… | 32 | 3090 | abandoned |
| NikolaiT/GoogleScraper GoogleScraper is a Python module and CLI tool for scraping search engine results from Google, Bing, Yandex, DuckDuckGo and others, with sup… | 32 | 2874 | abandoned |
| YahooArchive/anthelion Anthelion is an Apache Nutch plugin for focused crawling of semantic data embedded in HTML pages. It uses an online learning classifier to … | 10 | 2827 | abandoned |
| lanbing510/DouBanSpider A Python web scraper for Douban Books that crawls book listings by tag, storing ratings and review counts into Excel files. The author also… | 32 | 2786 | abandoned |
| loadchange/amemv-crawler A Python 3 script that downloads all videos from a specified Douyin (TikTok China) user account, as well as all videos under a given challe… | 32 | 2641 | abandoned |
| taspinar/twitterscraper A Python library that scrapes tweets and user information from Twitter using requests and BeautifulSoup, without relying on Twitter's offic… | 23 | 2461 | abandoned |
| scrapoxy/scrapoxy Scrapoxy was an open-source proxy manager for web scraping that aggregated proxies from cloud providers and other sources behind a single A… | 62 | 2414 | abandoned |
| egrcc/zhihu-python A Python 2.7 library for scraping content from Zhihu, a Chinese Q&A platform, including questions, answers, users, and favorites. It can ex… | 32 | 2335 | abandoned |
| chiphuyen/lazynlp A Python library for crawling, cleaning, and deduplicating web pages to build massive monolingual text datasets, suitable for training lang… | 23 | 2284 | abandoned |
| PaulMcInnis/JobFunnel JobFunnel is a Python CLI tool that scrapes job postings from multiple job websites (Indeed, Glassdoor, LinkedIn) into a single deduplicate… | 10 | 2180 | abandoned |
| minimaxir/facebook-page-post-scraper A Python script collection that scrapes all posts, reactions, and comments from public Facebook Pages and open Groups via the Facebook Grap… | 10 | 2135 | abandoned |
| simplecrawler/simplecrawler simplecrawler is a flexible, event-driven web crawler library for Node.js with a configurable queue system, robots.txt support, and link di… | 10 | 2134 | abandoned |
| althonos/InstaLooter InstaLooter is a Python CLI tool that downloads pictures and videos from Instagram profiles without using the official API. It is a re-impl… | 23 | 2098 | abandoned |
| cycz/jdBuyMask A Python automation tool that monitored JD.com (Jingdong) for face mask stock during the COVID-19 pandemic and automatically placed purchas… | 32 | 1849 | abandoned |
| hu17889/go_spider go_spider is a concurrent web crawler framework written in Go, designed for crawling vertical communities with a flexible, modular architec… | 23 | 1818 | abandoned |
| node-js-libs/node.io node.io is a Node.js web scraping and data extraction library originally written in 2010. It is explicitly no longer maintained, with the a… | 32 | 1792 | abandoned |
| bughandler/cnki-downloader A small desktop tool for searching and downloading academic literature from CNKI (China National Knowledge Infrastructure). Its backend int… | 48 | 1762 | abandoned |
| ZFC-Digital/puppeteer-real-browser A Node.js library that wraps Puppeteer with a real-browser profile to bypass bot detection systems like Cloudflare and Turnstile captchas. … | 35 | 1642 | abandoned |
| erma0/douyin A Python crawler for Douyin (Chinese TikTok) that collected public data such as account profiles, likes, favorites, music, hashtags, search… | 76 | 1602 | abandoned |
| fossasia/loklak_wok_android Loklak Wok is an Android app that acts as a harvesting peer for the loklak_server, collecting social media messages (tweets) and pushing th… | 10 | 1558 | abandoned |
| GravityLabs/goose Goose is a Scala library (originally Java) that extracts the main body text, metadata, publish date, embedded videos, and top image from ne… | 10 | 1526 | abandoned |
| yhat/scrape A Go library providing a higher-level interface over golang.org/x/net/html for web scraping. It offers generic tree traversal helpers like … | 10 | 1515 | abandoned |
| qinxuye/cola Cola is a high-level distributed crawling framework in Python for scraping pages and extracting structured data from websites. The same cra… | 10 | 1499 | abandoned |
| johntitus/node-horseman A Node.js library providing a chainable, Promise-based API for controlling the PhantomJS headless browser, supporting page navigation, form… | 32 | 1485 | abandoned |
| fossasia/loklak_scraper_js A collection of JavaScript scrapers for the loklak project that extract data from websites like Twitter and Quora and output JSON resemblin… | 10 | 1473 | abandoned |
| keenwon/antcolony AntColony is a Node.js-based BitTorrent DHT network crawler that collects active infohashes, downloads and parses torrent files, and stores… | 32 | 1456 | abandoned |
| jamesturk/scrapeghost scrapeghost is an experimental Python library that uses OpenAI's GPT models to scrape structured data from websites without writing page-sp… | 49 | 1442 | abandoned |
| OpnTec/parliament-scraper A collection of scrapers (in Python, Ruby, and Scala) that download public parliamentary data such as written questions from the EU Parliam… | 32 | 1403 | abandoned |
| leonardocardoso/SwiftLinkPreview A Swift library that generates link previews from URLs by extracting titles, relevant text, and images. It supports iOS, macOS, watchOS, an… | 10 | 1385 | abandoned |
| dinubs/jam-api Jam API is a hosted service (self-hostable Node.js app) that turns any website into a JSON API by extracting data using CSS query selectors… | 32 | 1363 | abandoned |
| Vespa314/bilibili-api A collection of Bilibili (B站) API documentation and Python tooling for scraping video, user, comment, danmaku, and bangumi data. It include… | 32 | 1357 | abandoned |
| Jefferson-Henrique/GetOldTweets-python A Python library that retrieves old tweets by mimicking the JSON calls Twitter Search makes in the browser, bypassing the official API's ti… | 32 | 1340 | abandoned |
| adyzng/jd-autobuy A Python 2.7 command-line scraper that logs into JD.com (via QR code or credentials), monitors product stock and price, and automatically p… | 32 | 1309 | abandoned |
| LiuRoy/zhihu_spider A Scrapy-based web crawler that scrapes Zhihu user profiles and their follower/following relationship graphs, storing data in MongoDB. It s… | 32 | 1282 | abandoned |