domain: crawlers
581 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| Firecrawl Firecrawl is an open-source web scraping and crawling API that turns websites into clean Markdown, structured JSON, screenshots, and other … | 88 | 172809 | active |
| unclecode/crawl4ai Crawl4AI is an open-source Python library that crawls websites with a headless browser and converts pages into clean, LLM-ready Markdown fo… | 89 | 79471 | active |
| D4Vinci/Scrapling Scrapling is an adaptive Python web scraping framework that handles everything from single requests to full-scale concurrent crawls. It fea… | 89 | 76660 | active |
| Panniantong/Agent-Reach Agent Reach is a Python CLI and MCP tool that gives AI agents the ability to read and search across platforms like Twitter, Reddit, YouTube… | 78 | 75643 | active |
| scrapy/scrapy Scrapy is a fast, high-level web crawling and scraping framework for Python used to extract structured data from websites. It provides a fu… | 99 | 64048 | stable |
| NanmiCoder/MediaCrawler MediaCrawler is a Python-based multi-platform social media crawler that scrapes notes, videos, posts, and comments from Xiaohongshu, Douyin… | 73 | 63826 | active |
| soimort/you-get You-Get is a tiny command-line utility written in Python to download media content (videos, audios, images) from the web when no other hand… | 67 | 56871 | active |
| NaiboWang/EasySpider EasySpider is a free, open-source visual no-code web crawler and browser automation (RPA) tool where users design scraping tasks by clickin… | 87 | 44442 | active |
| iawia002/lux Lux is a fast and simple video downloader written in Go, available both as a CLI tool and a library. It supports downloading videos from ma… | 56 | 31656 | active |
| CloakHQ/CloakBrowser CloakBrowser is a stealth Chromium distribution with 73 source-level C++ fingerprint patches that passes major bot-detection systems like C… | 81 | 30849 | active |
| feder-cr/Jobs_Applier_AI_Agent_AIHawk AIHawk is an open-source Python AI agent that automates job applications by scraping job postings via browser automation and auto-applying … | 58 | 30260 | active |
| ScrapeGraphAI/Scrapegraph-ai ScrapeGraphAI is a Python library that uses LLMs and direct graph logic to build web scraping pipelines for websites and local documents (X… | 88 | 29959 | active |
| ArchiveBox/ArchiveBox ArchiveBox is an open-source, self-hosted web archiving application that saves URLs, browser history, bookmarks, RSS feeds, and social medi… | 83 | 28188 | active |
| Crawlee Crawlee is a web scraping and browser automation library for Node.js/TypeScript (with a Python port) for building reliable crawlers. It int… | 98 | 25516 | stable |
| gocolly/colly Colly is a fast and elegant web scraping and crawling framework for Go. It provides a clean callback-based API for making HTTP requests, pa… | 66 | 25483 | stable |
| jhao104/proxy_pool A Python proxy pool service that periodically scrapes free proxies from 15+ sources, validates them, and stores them in Redis/SSDB. It expo… | 62 | 23642 | active |
| BuilderIO/gpt-crawler A Node.js/TypeScript crawler that scrapes one or more websites and generates knowledge files (JSON) for creating custom GPTs or OpenAI assi… | 31 | 22393 | active |
| h4ckf0r0day/obscura Obscura is an open-source headless browser engine written in Rust that runs JavaScript via V8 and speaks the Chrome DevTools Protocol, acti… | 81 | 22334 | active |
| xifangczy/cat-catch Cat-catch is an open-source browser extension that sniffs and lists media resources (video, audio, images) loaded by the current web page. … | 97 | 21548 | active |
| Evil0ctal/Douyin_TikTok_Download_API A high-performance asynchronous Python service that scrapes and parses data from Douyin, TikTok, Kuaishou, and Bilibili, exposing a FastAPI… | 55 | 19679 | active |
| mikf/gallery-dl gallery-dl is a command-line program to download image galleries and collections from many image hosting sites such as Pixiv, Danbooru, Dev… | 96 | 19332 | active |
| projectdiscovery/katana Katana is a fast, configurable web crawling and spidering framework written in Go, supporting both standard HTTP-based and headless browser… | 94 | 17348 | active |
| getmaxun/maxun Maxun is an open-source no-code platform for web scraping, crawling, search, and AI-powered data extraction that turns websites into struct… | 94 | 17299 | active |
| MODSetter/SurfSense SurfSense is an open-source, self-hostable NotebookLM alternative that combines a personal knowledge base with live open-web research conne… | 90 | 16017 | active |
| FlareSolverr/FlareSolverr FlareSolverr is a proxy server that bypasses Cloudflare and DDoS-GUARD protection by solving browser challenges with Selenium and undetecte… | 95 | 15316 | active |
| PuerkitoBio/goquery goquery is a Go library that provides a jQuery-like syntax for parsing, traversing, and manipulating HTML documents, built on Go's net/html… | 86 | 14981 | stable |
| browserless/browserless Browserless is a platform for deploying and managing headless browsers (Chrome, Firefox, WebKit) in Docker, usable self-hosted or via their… | 99 | 13634 | active |
| pystardust/ani-cli ani-cli is a POSIX shell script CLI tool for browsing and watching anime from the terminal, scraping the anidb site and playing videos via … | 97 | 13602 | active |
| instaloader/instaloader Instaloader is a Python command-line tool and library for downloading pictures, videos, captions, comments, and metadata from Instagram. It… | 87 | 13240 | active |
| ultrafunkamsterdam/undetected-chromedriver A Python library that patches Selenium's ChromeDriver binary so automated Chrome sessions avoid detection by anti-bot systems like Cloudfla… | 46 | 12808 | active |
| JoeanAmier/XHS-Downloader A tool for extracting links and downloading content (images, videos, live photos) from XiaoHongShu (RedNote) posts. It ships as a Python TU… | 76 | 12489 | active |
| g1879/DrissionPage DrissionPage is a Python-based web automation library that combines browser control (via Chromium's DevTools Protocol, without WebDriver) w… | 70 | 12405 | active |
| crawlab-team/crawlab Crawlab is a Go-based distributed web crawler management platform with a web UI for managing, scheduling, and monitoring spiders written in… | 52 | 12262 | active |
| jina-ai/reader Jina AI Reader converts any URL into LLM-friendly markdown via the r.jina.ai prefix, and searches the web into markdown via s.jina.ai. It r… | 62 | 11912 | active |
| code4craft/webmagic WebMagic is a scalable web crawler framework for Java covering the full crawl lifecycle: downloading, URL management, content extraction (X… | 60 | 11678 | active |
| daijro/camoufox Camoufox is an open-source anti-detect browser built on Firefox for web scraping and AI agents, with browser fingerprint spoofing and Playw… | 90 | 11455 | active |
| jhy/jsoup jsoup is a Java library for parsing, manipulating, and cleaning real-world HTML and XML, implementing the WHATWG HTML5 specification to pro… | 95 | 11387 | stable |
| ihmily/DouyinLiveRecorder A Python-based livestream recording application that can run unattended in a loop and record multiple streams simultaneously from 40+ platf… | 64 | 10786 | active |
| pinchtab/pinchtab PinchTab is a standalone Go HTTP server that gives AI agents direct control over Chrome via CDP, with stealth injection and multi-instance … | 82 | 10144 | active |
| dataabc/weiboSpider A Python command-line crawler that scrapes posts and profile data from Sina Weibo users, writing results to txt/csv/json files or MySQL/Mon… | 62 | 9695 | active |
| RSS-Bridge/rss-bridge RSS-Bridge is a PHP web application that generates RSS, Atom, and JSON feeds for websites that don't offer one, using hundreds of site-spec… | 75 | 9190 | active |
| kepano/defuddle Defuddle is a TypeScript library, CLI, and hosted service that extracts the main content from web pages, removing clutter like comments, si… | 87 | 9166 | active |
| qiye45/wechatDownload A desktop tool for batch downloading WeChat Official Account (公众号) articles, saving them as html/mhtml/md/pdf/docx/csv and preserving embed… | 90 | 9105 | active |
| Thysrael/Horizon Horizon is an AI-powered news aggregation platform that fetches content from sources like Hacker News, GitHub, RSS, Reddit, and Telegram, s… | 59 | 9040 | active |
| HerbertHe/iptv-sources A service that automatically aggregates and updates IPTV m3u playlist sources from multiple public repositories, with EPG data and optional… | 71 | 8921 | active |
| jo-inc/camofox-browser A self-hosted anti-detection browser server for AI agents, wrapping the Camoufox Firefox fork that spoofs fingerprints at the C++ level. It… | 82 | 8894 | active |
| eze-is/web-access An Agent Skill that gives AI coding agents (Claude Code, Cursor, Gemini CLI, etc.) full web access capabilities: three-tier channel dispatc… | 65 | 8745 | active |
| hardikvasa/google-images-download A Python command-line tool that searches and downloads hundreds of images from Google Images to local storage. It uses Selenium with Chrome… | 70 | 8684 | active |
| kangvcar/InfoSpider InfoSpider is an open-source Python toolbox that crawls a user's own personal data from dozens of Chinese and international services (email… | 58 | 8247 | active |
| alirezamika/autoscraper AutoScraper is a Python library that automatically learns scraping rules from a URL or HTML content plus a list of sample data you want to … | 66 | 7904 | stable |
| andeya/pholcus Pholcus is a distributed, high-concurrency web crawler framework written in pure Go. It supports standalone, server, and client modes with … | 90 | 7577 | active |
| mgdm/htmlq htmlq is a command-line tool, like jq but for HTML, that extracts content from HTML documents using CSS selectors. Written in Rust, it read… | 61 | 7576 | stable |
| Steel Browser Steel Browser is an open-source browser API and sandbox that manages browser sessions, proxies, stealth, and lifecycle so developers can bu… | 89 | 7545 | active |
| cv-cat/Spider_XHS A Python library that reverse-engineers Xiaohongshu (Little Red Book) signature algorithms and wraps the platform's PC, creator, and Pugong… | 88 | 7423 | active |
| berstend/puppeteer-extra A modular plugin framework for Puppeteer (and Playwright via playwright-extra) that extends headless browser automation with drop-in plugin… | 32 | 7398 | active |
| autoscrape-labs/pydoll Pydoll is a Python library for automating Chromium-based browsers directly over the Chrome DevTools Protocol, with no WebDriver binary and … | 85 | 7048 | active |
| hect0x7/JMComic-Crawler-Python A Python library providing an API client for the JMComic (18comic) site, supporting both web and mobile endpoints, with album downloading, … | 97 | 6997 | stable |
| mishushakov/llm-scraper A TypeScript library that turns any webpage into structured data using LLMs, built on Playwright and the Vercel AI SDK. It supports multipl… | 67 | 6915 | active |
| bda-research/node-crawler node-crawler is a TypeScript web crawler/spider library for Node.js that fetches pages and provides server-side DOM parsing with automatic … | 91 | 6799 | active |
| VeNoMouS/cloudscraper A Python library that wraps Requests to bypass Cloudflare's anti-bot protection pages (IUAM), supporting challenge types v1, v2, v3, and Tu… | 34 | 6726 | active |
| adbar/trafilatura Trafilatura is a Python package and command-line tool for crawling the web and extracting main text, metadata, and comments from raw HTML w… | 93 | 6709 | active |
| davidteather/TikTok-Api An unofficial Python API wrapper for TikTok.com that retrieves trending content, user information, hashtags, and video data without authent… | 92 | 6593 | active |
| brightdata/cli The official Bright Data CLI (npm package @brightdata/cli) providing terminal access to Bright Data's web scraping, search, and structured … | 81 | 6406 | active |
| lexiforest/curl_cffi curl_cffi is a Python binding for a curl-impersonate fork via cffi, providing an HTTP client that can impersonate browser TLS/JA3, HTTP/2, … | 99 | 6390 | active |
| chyroc/WechatSogou A Python library providing a scraping API for WeChat official accounts based on Sogou WeChat search. It lets you look up account info and r… | 54 | 6377 | active |
| Python3WebSpider/ProxyPool A self-hosted proxy pool service that periodically scrapes free proxy sites, stores and scores them in Redis, tests their availability, and… | 63 | 6243 | active |
| epiral/bb-browser bb-browser is a TypeScript CLI and MCP server that lets AI agents control your real Chrome browser using your existing login state, exposin… | 70 | 6128 | active |
| MontFerret/ferret Ferret is a declarative, expression-oriented query language (FQL) with an embeddable Go runtime for querying, transforming, and automating … | 83 | 6008 | active |
| hect0x7/JMComic-APK A GitHub Actions-based automation that periodically checks for updates to the JM Comic (18comic) Android APK and publishes new versions as … | 93 | 5992 | active |
| omkarcloud/botasaurus Botasaurus is an all-in-one Python web scraping framework with built-in anti-detection, caching, parallelization, and proxy support. It let… | 72 | 5688 | active |
| rmax/scrapy-redis Redis-based components for Scrapy that enable distributed crawling and scraping by sharing a Redis queue across multiple spider instances. … | 60 | 5642 | active |
| gosom/google-maps-scraper An open-source Go tool that scrapes Google Maps to extract business data such as names, addresses, phone numbers, websites, ratings, review… | 92 | 5630 | active |
| xuejianxianzun/PixivBatchDownloader A browser extension (Chrome, Edge, Firefox) for batch downloading illustrations, manga, Ugoira animations, and novels from Pixiv. It offers… | 98 | 5564 | active |
| browser-act/skills BrowserAct Skills is a Python-based browser automation CLI designed for AI agents, providing real-browser control with anti-bot evasion (st… | 60 | 5446 | active |
| AhmadIbrahiim/Website-downloader A Node.js web application that downloads the complete source code of any website, including all assets like JavaScripts, stylesheets, and i… | 76 | 5245 | active |
| Yuukiy/JavSP JavSP is a Python command-line tool that scrapes adult video (JAV) metadata from multiple websites, aggregates the data, and generates NFO … | 25 | 5136 | active |
| lc/gau gau (getallurls) is a Go CLI tool that fetches known URLs for a given domain from AlienVault's Open Threat Exchange, the Wayback Machine, C… | 56 | 5076 | active |
| apify/apify-mcp-server The Apify MCP Server exposes thousands of Apify Store scrapers, crawlers, and automation tools to AI agents via the Model Context Protocol,… | 84 | 5059 | active |
| jaypyles/Scraperr Scraperr is a self-hosted web scraping application with a web UI that lets users scrape websites without writing code, using XPath-based ex… | 10 | 4910 | active |
| MechanicalSoup/MechanicalSoup A Python library for automating interaction with websites, built on Requests and BeautifulSoup. It handles cookies, redirects, link followi… | 66 | 4888 | active |
| bjesus/pipet Pipet is a command-line web scraper written in Go that extracts data from online assets using HTML parsing, JSON parsing, and client-side J… | 24 | 4772 | active |
| l0o0/translators_CN A community-maintained collection of Zotero translators for Chinese academic and general websites, enabling Zotero to scrape citation metad… | 76 | 4721 | active |
| DedSecInside/TorBot TorBot is a Python CLI tool for OSINT on the dark web, crawling .onion sites over the Tor network and building link trees. It can save craw… | 97 | 4716 | active |
| xroche/httrack HTTrack is a free offline browser utility that recursively downloads websites to a local directory, rewriting links so the mirrored copy ca… | 99 | 4702 | stable |
| ultrafunkamsterdam/nodriver Nodriver is a fully asynchronous Python browser automation and web scraping library, and the official successor to Undetected-Chromedriver.… | 62 | 4699 | active |
| lecepin/WeChatVideoDownloader A convenient desktop GUI application for downloading videos from WeChat Channels (WeChat Video Accounts). It intercepts and captures video … | 10 | 4677 | active |
| d60/twikit Twikit is a free Python library that wraps Twitter's internal API, allowing posting, searching, and scraping tweets without an official API… | 60 | 4631 | active |
| dataabc/weibo-crawler A Python crawler for Sina Weibo that scrapes user profiles and posts, exporting data to CSV, JSON, MySQL, MongoDB, or SQLite, and optionall… | 74 | 4625 | active |
| Keiyoushi Extensions A community-maintained repository of extensions (APKs) for Mihon and its forks, providing manga source plugins. The source code for the ext… | 71 | 4595 | active |
| joeyism/linkedin_scraper A Python library that scrapes LinkedIn for user, company, and job data using Playwright with an async API. It provides Pydantic data models… | 82 | 4452 | active |
| sparklemotion/mechanize Mechanize is a Ruby library for automating interaction with websites. It handles cookies, redirects, link following, and form submission wh… | 91 | 4439 | stable |
| rachelos/we-mp-rss A self-hosted WeChat official account (公众号) subscription assistant that scrapes articles, generates RSS feeds, and converts content to Mark… | 84 | 4396 | active |
| UltimaHoarder/UltimaScraper A Python-based scraper that downloads all media (photos, videos) from OnlyFans accounts using the user's own session authentication. It sto… | 23 | 4271 | active |
| kanasimi/work_crawler A multi-language downloader application that batch-downloads web novels (converting them to EPUB) and comics from a large list of Chinese, … | 66 | 4193 | active |
| Patchright Patchright is a patched, undetected fork of the Playwright browser automation framework that evades bot-detection systems like Cloudflare. … | 93 | 4191 | active |
| speedyapply/JobSpy JobSpy is a Python library that scrapes job postings from popular job boards like LinkedIn, Indeed, Glassdoor, Google, and ZipRecruiter con… | 57 | 4165 | active |
| dotnetcore/DotnetSpider DotnetSpider is a .NET Standard web crawling and scraping framework that is lightweight, efficient, and cross-platform. It supports distrib… | 57 | 4138 | active |
| Lucksi/Mr.Holmes Mr.Holmes is a Python-based OSINT (open-source intelligence) CLI tool that gathers information about usernames, domains, phone numbers, and… | 53 | 4112 | active |
| nghuyong/WeiboSpider A continuously maintained Python web scraping tool for Sina Weibo built on Scrapy and the new weibo.com API. It collects user profiles, pos… | 73 | 4109 | active |
| RipMeApp/ripme RipMe is a cross-platform Java application that bulk-downloads image albums from websites like Reddit, Imgur, Twitter, Instagram, and Tumbl… | 98 | 4104 | active |
page 1 / 6 next →