domain: crawlers
581 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| watercrawl/WaterCrawl WaterCrawl is a self-hostable web application (Python/Django/Scrapy/Celery) that crawls websites and transforms web content into LLM-ready … | 82 | 2010 | active |
| supermemoryai/markdowner Markdowner is a fast web service that converts any website into LLM-ready markdown, with optional LLM filtering, detailed responses, and au… | 24 | 1999 | active |
| nottelabs/notte Notte is a full-stack framework and cloud platform for building, deploying, and scaling AI web agents and browser automations. It combines … | 84 | 1997 | active |
| zhegexiaohuozi/SeimiCrawler SeimiCrawler is an agile, standalone, distributed Java crawler framework inspired by Python's Scrapy, with deep Spring Boot integration and… | 85 | 1990 | active |
| AWeirdDev/flights fast-flights is a Python library that scrapes Google Flights by generating Base64-encoded Protobuf query strings, returning strongly-typed … | 86 | 1939 | active |
| feder-cr/invisible_playwright A Python library that provides an antidetect, stealth-patched Firefox build for Playwright, with fingerprints set at the C++ engine level a… | 81 | 1938 | active |
| damoeb/rss-proxy RSS-proxy is a self-hostable web service that generates RSS, ATOM, or JSON feeds from almost any static website by analyzing its HTML struc… | 32 | 1924 | active |
| MarginaliaSearch/MarginaliaSearch Marginalia Search is an independent, open-source internet search engine that indexes text-oriented, non-commercial, small and old websites.… | 65 | 1923 | active |
| KEV0143/Parser-Chitai-Gorod A Python-based scraper for the Russian online bookstore Chitai-Gorod that collects book URLs across catalog pages and extracts structured p… | 29 | 1914 | active |
| extractus/article-extractor A TypeScript library that extracts the main article content, title, image, and metadata from a given URL or raw HTML string. It supports cu… | 98 | 1909 | active |
| trevorhobenshield/twitter-api-client A Python library implementing X/Twitter's v1, v2, and GraphQL APIs for automation and scraping. It supports account actions like tweeting, … | 30 | 1894 | active |
| 404-novel-project/novel-downloader An extensible userscript (Tampermonkey/Greasemonkey/Violentmonkey) that downloads novels from many Chinese web novel sites and exports them… | 76 | 1877 | active |
| coder-hxl/x-crawl x-crawl is a flexible Node.js crawler library that supports crawling dynamic pages, static pages, API data, and files, with optional AI ass… | 66 | 1877 | active |
| ThePhaseless/Byparr Byparr is a self-hosted Python service that solves antibot browser challenges (like Cloudflare checks) and returns valid clearance cookies … | 90 | 1865 | active |
| sarperavci/GoogleRecaptchaBypass A Python library that automatically solves Google reCAPTCHA v2 challenges in under five seconds using browser automation with DrissionPage … | 68 | 1855 | active |
| microlinkhq/browserless A Node.js library that wraps Puppeteer to provide a production-ready headless Chrome/Chromium driver with built-in screenshot, PDF generati… | 95 | 1831 | active |
| tryolabs/requestium Requestium is a Python library that merges Requests, Selenium, and Parsel into a single integrated tool for web automation. It lets scripts… | 77 | 1830 | active |
| enetx/surf Surf is an advanced HTTP client library for Go with fluent, chainable API design. It supports browser impersonation (Chrome/Firefox), JA3/J… | 83 | 1808 | active |
| deweizhu/bookget bookget is a Go-based command-line tool for downloading digitized ancient books and rare texts from 50+ digital libraries. It ships prebuil… | 60 | 1764 | active |
| LoseNine/ruyipage RuyiPage is a Python browser automation framework built on Firefox and the WebDriver BiDi protocol, shipping with an anti-detection Firefox… | 79 | 1759 | active |
| josh0xA/darkdump Darkdump is an open-source OSINT tool for querying multiple dark web search engines and scraping onion site results for emails, metadata, k… | 70 | 1757 | active |
| website-scraper/node-website-scraper A Node.js library that downloads entire websites to a local directory, including HTML, CSS, images, and JavaScript assets. It parses HTTP r… | 72 | 1751 | active |
| egoist/sitefetch A Node.js CLI tool that crawls an entire website and saves its pages as a single text file, using Mozilla Readability to extract clean cont… | 22 | 1736 | active |
| claffin/cloudproxy CloudProxy is a self-hosted Python tool that provisions and manages proxy servers across multiple cloud providers, rotating IPs to improve … | 81 | 1722 | active |
| Jules-WinnfieldX/CyberDropDownloader A Python-based bulk downloader that scrapes and downloads files from Cyberdrop.me and dozens of other file hosts and image galleries. It is… | 10 | 1720 | active |
| 3441293738/creatorhub CreatorHub is a self-hosted web panel built with Python and FastAPI for managing, monitoring, scraping, downloading, and publishing content… | 58 | 1712 | active |
| Python3Spiders/WeiboSuperSpider A Weibo (Chinese microblog) scraping toolbox in Python covering users, topics, and comments, with extras like image downloading, sentiment … | 75 | 1705 | active |
| oxylabs/google-play-scraper A free Python-based Google Play Store scraper that collects public app, movie, and book data via search queries. It is a companion tool to … | 67 | 1693 | active |
| vibheksoni/stealth-browser-mcp A Python MCP server that exposes stealth browser automation (via nodriver and Chrome DevTools Protocol) to AI agents, letting them navigate… | 61 | 1674 | active |
| rushter/selectolax Selectolax is a fast Python HTML5 parser library written in Cython, binding to the Modest and Lexbor parsing engines. It provides CSS selec… | 94 | 1665 | active |
| MgArcher/Text_select_captcha A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio… | 69 | 1656 | active |
| srx-2000/spider_collection A collection of Python web crawler scripts targeting sites like Bilibili, Zhihu, Weibo, NetEase Music, GitHub, and Anjuke, built with reque… | 32 | 1646 | active |
| Jesseovo/last30days-skill-cn An AI Agent skill (for Claude Code / OpenClaw) that automatically searches content from the last 30 days across 8 major Chinese internet pl… | 56 | 1642 | active |
| wu529778790/panhub.shenzjd.com PanHub is a self-hostable netdisk search aggregator that combines results from Quark, Aliyun Drive, Baidu Netdisk, 115, Thunder and 80+ Tel… | 62 | 1617 | active |
| hartator/wayback-machine-downloader A Ruby command-line tool that downloads an entire website from the Internet Archive Wayback Machine, restoring original files and directory… | 23 | 5930 | maintenance |
| ArchiveTeam/grab-site grab-site is a preconfigured web crawler for archiving websites, producing WARC files via a fork of wpull. It includes a dashboard for moni… | 43 | 1607 | active |
| nerevu/riko riko is a pure Python stream processing library modeled after Yahoo! Pipes, combining reusable, configuration-driven modular pipes with syn… | 88 | 1606 | active |
| matthewmueller/x-ray x-ray is a Node.js web scraping library that lets you define flexible schemas to structure data from any website using jQuery-like selector… | 66 | 5907 | maintenance |
| Altimis/Scweet Scweet is a Python library and CLI for scraping tweets, profile timelines, followers, following lists, and user profiles from Twitter/X wit… | 78 | 1604 | active |
| lanyeeee/jmcomic-downloader A multi-threaded GUI downloader for the 18comic.vip (jmcomic) manga site, built with Tauri (Rust backend, Vue frontend). It supports search… | 89 | 1587 | active |
| s0md3v/uro uro is a Python CLI tool that declutters URL lists for crawling and security testing without making any HTTP requests. It removes duplicate… | 29 | 1587 | stable |
| xnl-h4ck3r/xnLinkFinder xnLinkFinder is a Python CLI tool that discovers endpoints, potential parameters, target-specific wordlists, and secrets for a given target… | 78 | 1585 | active |
| ttttmr/Wechat2RSS Wechat2RSS is a service and self-hostable tool that converts WeChat official account (公众号) articles into RSS feeds, aiming for updates with… | 75 | 1557 | active |
| ulixee/hero Hero is a headless web browser built specifically for web scraping, powered by Chrome and controlled from NodeJS with a fully compliant DOM… | 69 | 1554 | active |
| Rhizome-Conifer/conifer Conifer is an open-source web archiving platform for capturing, replaying, and sharing collections of archived web pages through a user-fri… | 74 | 1549 | active |
| skernelx/tavily-key-generator A Python toolkit that automates signup flows for Tavily, Firecrawl, and Exa using real browser automation (Playwright/Camoufox), Turnstile … | 48 | 1549 | active |
| yujiosaka/headless-chrome-crawler A Node.js library providing a distributed web crawler powered by Headless Chrome via Puppeteer. It can crawl JavaScript-rendered (SPA) webs… | 23 | 5635 | maintenance |
| oxylabs/ai-map-py AI-Map is a Python SDK client for Oxylabs AI Studio's AI-powered website mapping service, which discovers and extracts relevant URLs from a… | 50 | 1537 | active |
| LifeActor/ykdl YouKuDownLoader (ykdl) is a Python command-line video downloader focused on China mainland video sites, forked from you-get with restructur… | 47 | 1533 | active |
| justfoolingaround/animdl animdl is a lightweight Python CLI tool that scrapes, streams, and downloads anime episodes from supported providers. It supports quality s… | 32 | 1522 | active |
| tidyverse/rvest rvest is an R package from the tidyverse for scraping (harvesting) data from web pages, inspired by Beautiful Soup and RoboBrowser. It prov… | 43 | 1520 | active |
| Danny-Dasilva/CycleTLS CycleTLS is a Go library with a JavaScript/TypeScript wrapper that lets clients spoof TLS/JA3 (and JA4) fingerprints when making HTTP reque… | 73 | 1518 | active |
| SpiderClub/haipproxy A high-availability distributed IP proxy pool built with Scrapy and Redis that scrapes free proxies from the internet, validates them, and … | 23 | 5523 | maintenance |
| oxylabs/browser-agent-py A Python SDK for Oxylabs AI Studio's Browser Agent, a cloud service that automates real-user browsing tasks (clicking, typing, scrolling, s… | 51 | 1505 | active |
| requests-cache/requests-cache requests-cache is a persistent HTTP caching library for Python's requests library, providing a drop-in CachedSession and optional global pa… | 77 | 1501 | stable |
| JustAnotherArchivist/snscrape snscrape is a Python-based scraper for social networking services that extracts posts, profiles, hashtags, and search results from platform… | 32 | 5444 | maintenance |
| roach-php/core Roach is a complete web scraping and crawling toolkit for PHP, heavily inspired by Python's Scrapy. It lets developers define spiders that … | 45 | 1455 | active |
| submato/xhscrawl A Python-based reverse-engineering toolkit for Xiaohongshu (XHS) web APIs, focusing on generating the encrypted x-s signature parameter via… | 72 | 1452 | active |
| tinyfish-io/agentql AgentQL is a suite of tools for extracting structured data and automating workflows on live websites using an AI-powered natural language q… | 70 | 1451 | active |
| scrapy-plugins/scrapy-playwright A Scrapy download handler that uses Playwright for Python to fetch pages, enabling scraping of JavaScript-rendered sites while keeping the … | 86 | 1439 | active |
| ShilongLee/Crawler A self-hostable crawler API server that exposes HTTP endpoints for scraping public data from Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo… | 53 | 1435 | active |
| drawrowfly/tiktok-scraper A TypeScript library and CLI tool that scrapes TikTok metadata from user, hashtag, trend, and music pages and downloads video posts without… | 23 | 5173 | maintenance |
| zhaoolee/garss Garss (嘎!RSS) is a self-hosted RSS aggregation and reading system that uses GitHub Actions to collect hundreds of RSS feeds and render them… | 82 | 1426 | active |
| rebrowser/rebrowser-patches A collection of source-code patches for Puppeteer and Playwright that fix automation leaks and help avoid bot detection systems like Cloudf… | 34 | 1424 | active |
| cdpdriver/zendriver Zendriver is an async-first Python web scraping and browser automation framework built on the Chrome Devtools Protocol, forked from nodrive… | 88 | 1409 | active |
| KoalaBear84/OpenDirectoryDownloader A cross-platform C#/.NET command-line tool that indexes open directory listings across 130+ supported formats, including FTP(S), Google Dri… | 94 | 1389 | active |
| lorey/mlscraper mlscraper is a Python library that automatically extracts structured data from HTML pages using machine learning. Instead of writing CSS se… | 23 | 1385 | active |
| Avnsx/fansly-downloader A Python-based tool for bulk downloading photos, videos, and audio from fansly.com, also shipped as a standalone Windows executable. It sup… | 10 | 1384 | active |
| CIRCL/AIL-framework AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T… | 67 | 1378 | active |
| mattsse/chromiumoxide chromiumoxide is a Rust library providing a high-level async API for controlling Chrome or Chromium via the Chrome DevTools Protocol. It ca… | 75 | 1375 | active |
| okfn-brasil/querido-diario Querido Diário is an open-source project by Open Knowledge Brasil that scrapes and aggregates Brazilian municipal official gazettes (diário… | 75 | 1373 | active |
| xiaohucode/xiangse A curated collection of video and manga source plugins (.xbs files) for the Xiangse Guige (香色闺阁) reading/media app, imported via URL. It ag… | 10 | 1359 | active |
| scrapy/parsel Parsel is a BSD-licensed Python library for extracting data from HTML, XML, and JSON documents using CSS selectors, XPath expressions, JMES… | 88 | 1352 | active |
| raznem/parsera Parsera is a lightweight Python library for scraping websites using LLMs, letting users define elements to extract with natural-language de… | 54 | 1350 | active |
| LeetaoGoooo/RSSAid RSSAid is a Flutter-based mobile app that complements RSSHub by helping users discover and subscribe to RSS feeds from websites, similar to… | 76 | 1345 | active |
| philippta/flyscrape Flyscrape is a standalone command-line web scraping tool written in Go that lets users write extraction logic in JavaScript with a jQuery-l… | 42 | 1345 | active |
| SpiderClub/weibospider A distributed web crawler for Sina Weibo (Chinese microblogging platform) built with Python, Celery, and requests. It scrapes user profiles… | 32 | 4793 | maintenance |
| mvdbos/php-spider A configurable and extensible PHP web spider library for crawling websites. It supports breadth-first and depth-first traversal, URI discov… | 75 | 1341 | active |
| tophubs/TopList TopList (今日热榜) is a self-hosted aggregation website that collects trending headlines from popular sites like Zhihu, Hupu, and V2EX. It is w… | 32 | 4730 | maintenance |
| karust/openserp OpenSERP is a self-hosted, MIT-licensed SERP API and CLI written in Go that returns structured search results from Google, Bing, Yandex, Ba… | 91 | 1306 | active |
| yasserg/crawler4j crawler4j is an open-source web crawler library for Java that provides a simple interface for building multi-threaded web crawlers in minut… | 23 | 4618 | maintenance |
| dwisiswant0/go-dork go-dork is a fast command-line dork scanner written in Go that automates Google dorking across multiple search engines. It supports Google,… | 23 | 1301 | stable |
| bookstairs/bookhunter bookhunter is a Go command-line tool for scraping and downloading ebooks from sources like Talebook, SoBooks, Telegram channels, and China'… | 54 | 1295 | active |
| techwithtim/Price-Tracking-Web-Scraper A full-stack price tracking application that scrapes product prices (currently Amazon.ca) using Playwright and Bright Data's Scraping Brows… | 29 | 1294 | active |
| zohaibbashir/Google-Maps-Scrapper A Python CLI script built on Playwright that scrapes Google Maps listings to extract business details such as name, address, website, phone… | 65 | 1289 | active |
| oxylabs/paid-proxy-servers A promotional GitHub repository for Oxylabs' commercial paid proxy services, covering residential, mobile, datacenter, ISP, and SOCKS5 prox… | 59 | 1289 | active |
| TheBeastLT/torrentio-scraper Torrentio is a Stremio addon ecosystem that scrapes public torrent providers and serves the results as Stremio stream results. The reposito… | 77 | 1280 | active |
| sardanioss/httpcloak httpcloak is a Go HTTP client library that reproduces browser-identical TLS, HTTP/2, and HTTP/3 fingerprints (JA3/JA4, Akamai, header order… | 61 | 1276 | active |
| kkangert/kspider Kspider is a self-hosted visual web scraping platform written in Java where users define crawler workflows as flowcharts without writing ba… | 14 | 1269 | active |
| minsight-ai-info/AI-Search-Hub AI Search Hub is an open-source Skill that aggregates native AI search capabilities from platforms like Gemini, Grok, Doubao, and Yuanbao i… | 50 | 1250 | active |
| egbertbouman/youtube-comment-downloader A Python script and library for downloading YouTube video comments without using the official YouTube API. It outputs comments in JSONL, JS… | 88 | 1247 | active |
| Decodo/Decodo Decodo (formerly Smartproxy) is a commercial rotating proxy network and web scraping platform offering 125M+ residential, mobile, ISP, and … | 69 | 1238 | active |
| raawaa/jav-scrapy A TypeScript-based Node.js CLI tool that batch-scrapes JAV (adult video) metadata, magnet links, and cover images from source websites. It … | 94 | 1235 | active |
| eatmoreduck/boss-zhipin-scraper A Python CLI scraper for BOSS Zhipin (zhipin.com) that connects to a locally logged-in Chrome via the Chrome DevTools Protocol to call the … | 57 | 1231 | active |
| iszhouhua/social-media-copilot An open-source browser extension (built with WXT and TypeScript) that scrapes data from Chinese social media platforms including Xiaohongsh… | 65 | 1228 | active |
| daijro/browserforge BrowserForge is a Python library that generates realistic browser headers and fingerprints, mimicking real-world browser, OS, and device di… | 56 | 1224 | active |
| firecrawl/web-agent An open-source TypeScript framework for building autonomous web research agents, layered from a Next.js chat template down to an agent core… | 50 | 1222 | active |
| cnbattle/douyin A Go-based crawler that scrapes Douyin (TikTok China) recommendation and search page video lists by controlling the mobile app on a real de… | 61 | 1219 | active |
| Silent1566/OmniBox-Spider A collection of spider (scraper) sources and interfaces for the OmniBox media application, aggregated from publicly available internet info… | 58 | 1213 | active |
| lexmount/moli Moli is a lightweight, fast headless browser built in Rust (on Servo technology) designed for AI agents to fetch, render, and extract web p… | 79 | 1206 | active |