domain: crawlers
581 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| IonicaBizau/scrape-it scrape-it is a Node.js web scraping library with a simple, declarative API for extracting data from HTML pages, built on top of tinyreq and… | 93 | 4073 | active |
| fake-useragent/fake-useragent A Python library that generates realistic, up-to-date browser user-agent strings from a bundled real-world database. It supports random or … | 10 | 4049 | active |
| MikeChongCan/scylla Scylla is an intelligent proxy pool service that automatically crawls, validates, and serves proxy IPs via a JSON API and web UI. It integr… | 34 | 4020 | active |
| binux/pyspider pyspider is a powerful web crawler (spider) system written in Python with a built-in WebUI for script editing, task monitoring, project man… | 10 | 16774 | maintenance |
| kingname/GeneralNewsExtractor GNE (GeneralNewsExtractor) is a Python library that extracts article title, author, publish time, body text, and images from news page HTML… | 81 | 3788 | active |
| edoardottt/cariddi Cariddi is a fast command-line web crawler written in Go that takes a list of domains, crawls URLs, and scans for endpoints, secrets, API k… | 84 | 3753 | active |
| Boris-code/feapder feapder is a powerful, easy-to-use Python web crawling framework offering four spider types (AirSpider, Spider, TaskSpider, BatchSpider) fo… | 74 | 3735 | active |
| JoeanAmier/TikTokDownloader DouK-Downloader (formerly TikTokDownloader) is an open-source Python tool for downloading and scraping content from Douyin and TikTok, incl… | 71 | 15584 | maintenance |
| oxylabs/google-ai-mode-scraper A code repository of examples for Oxylabs' Google AI Mode Scraper, a commercial API that sends prompts to Google AI Mode and returns parsed… | 61 | 3581 | active |
| codelucas/newspaper newspaper3k is a Python 3 library for discovering, downloading, and parsing news articles from websites. It extracts full text, authors, pu… | 66 | 15144 | maintenance |
| elliotgao2/toapi A Python library that turns any website into a JSON API by declaring fields with CSS/XPath selectors. It fetches and parses pages on demand… | 75 | 3555 | active |
| thomasdondorf/puppeteer-cluster A Node.js library that manages a pool of Puppeteer-controlled Chromium instances for running browser tasks in parallel. It handles job queu… | 60 | 3514 | stable |
| Gerapy/Gerapy Gerapy is a distributed crawler management framework built on Scrapy, Scrapyd, Django, and Vue.js. It provides a web dashboard for managing… | 63 | 3512 | active |
| google/robotstxt Google's production robots.txt parser and matcher, released as a C++ library (C++14 compliant). It implements the Robots Exclusion Protocol… | 67 | 3471 | stable |
| oxylabs/free-proxy-list A repository promoting Oxylabs' free tier of US datacenter proxies (HTTP/HTTPS/SOCKS5), offering 5 US IPs, 20 concurrent sessions, and 5GB … | 54 | 3434 | active |
| any4ai/AnyCrawl AnyCrawl is a Node.js/TypeScript web crawler and scraping service that converts websites into LLM-ready markdown/JSON data and extracts str… | 85 | 3415 | active |
| opsdisk/pagodo pagodo is a Python CLI tool that automates passive Google dork searches by scraping the Google Hacking Database (GHDB) and running those qu… | 49 | 3387 | active |
| white0dew/XiaohongshuSkills A Python CLI tool and agent Skill that automates Xiaohongshu (RED/RedNote) via Chrome DevTools Protocol, supporting publishing posts with i… | 60 | 3366 | active |
| oxylabs/oxylabs-ai-studio-py A Python SDK for Oxylabs AI Studio, a commercial API offering AI-powered web scraping, crawling, search, site mapping, and browser automati… | 59 | 3354 | active |
| oxylabs/google-news-scraper A free Python-based command-line tool from Oxylabs that scrapes Google News articles by topic, exporting headlines, URLs, and publication d… | 64 | 3329 | active |
| postaddictme/instagram-php-scraper A PHP library that scrapes Instagram via its web version to fetch account information, photos, videos, stories, and comments, with optional… | 33 | 3327 | active |
| oxylabs/amazon-scraper A Python-based free tool and sample code for Oxylabs' Amazon Scraper API that extracts product, search, offer listing, reviews, Q&A, best s… | 71 | 3326 | active |
| tamnd/kage kage is a Go CLI tool that clones websites into browsable offline folders by rendering each page in headless Chrome, snapshotting the final… | 79 | 3322 | active |
| psf/requests-html A Python library that combines HTTP requests with intuitive HTML parsing, offering CSS selector and XPath support plus JavaScript rendering… | 23 | 13814 | maintenance |
| internetarchive/heritrix3 Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler written in Java. It crawls websites and… | 95 | 3305 | active |
| oxylabs/chatgpt-scraper Code examples and documentation for Oxylabs' ChatGPT Scraper, a commercial Web Scraper API endpoint that sends prompts to ChatGPT and retur… | 62 | 3301 | active |
| php-curl-class/php-curl-class PHP Curl Class is a PHP library that wraps PHP's cURL extension in a simple object-oriented API for sending HTTP requests. It supports GET,… | 93 | 3297 | active |
| aapatre/Automatic-Udemy-Course-Enroller-GET-PAID-UDEMY-COURSES-for-FREE A Python script that scrapes coupon sites for free Udemy course coupons and automatically enrolls the user in paid courses for free using S… | 58 | 3291 | active |
| apache/nutch Apache Nutch is a highly extensible and scalable open-source web crawler built on Apache Hadoop data structures. It supports batch crawling… | 77 | 3277 | stable |
| itsOwen/CyberScraper-2077 CyberScraper 2077 is an AI-powered web scraping application with a Streamlit GUI that uses OpenAI, Gemini, or local Ollama models to intell… | 67 | 3245 | active |
| jsvine/waybackpack Waybackpack is a Python command-line tool that downloads the entire Wayback Machine archive for a given URL, saving every archived snapshot… | 30 | 3228 | active |
| oxylabs/ai-crawler-py A Python client library for Oxylabs AI Studio's AI-Crawler, a commercial service that crawls websites from a starting URL, uses natural lan… | 61 | 3218 | active |
| s0md3v/Photon Photon is a fast Python-based web crawler designed for OSINT (open-source intelligence) tasks. It extracts URLs, emails, social media accou… | 66 | 13146 | maintenance |
| devanshbatham/ParamSpider ParamSpider is a Python CLI tool that mines URLs for a domain or list of domains from web archives (Wayback Machine), filtering out uninter… | 55 | 3160 | active |
| omkarcloud/google-maps-scraper A desktop application and API for scraping Google Maps business data, extracting 50+ data points including emails, phone numbers, social pr… | 72 | 3130 | active |
| scrapy/scrapyd Scrapyd is a service daemon for deploying and running Scrapy spiders. It lets you upload Scrapy projects and control spiders via a JSON HTT… | 66 | 3101 | stable |
| symfony/panther Symfony Panther is a PHP library for browser testing and web scraping that drives real browsers (Chrome, Firefox) via the W3C WebDriver pro… | 73 | 3067 | active |
| Qianlitp/crawlergo crawlergo is a Go-based browser crawler that uses headless Chrome to discover URLs for web vulnerability scanners. It renders pages, fills … | 27 | 3034 | active |
| adryfish/fingerprint-chromium A fingerprint browser built on Ungoogled Chromium that lets users spoof or control browser fingerprint characteristics to avoid detection. … | 76 | 2991 | active |
| 5ime/video_spider A PHP-based web service that parses short-video links from platforms like Douyin, Kuaishou, Weibo, and Pipixia to return watermark-free vid… | 57 | 2975 | active |
| nashsu/AutoCLI AutoCLI is a blazing-fast, memory-safe command-line tool written in Rust that fetches information from 55+ websites (Twitter/X, Reddit, You… | 65 | 2950 | active |
| rust-headless-chrome/rust-headless-chrome A Rust library providing a high-level API to control headless Chrome or Chromium via the DevTools Protocol, serving as the Rust equivalent … | 81 | 2947 | active |
| striver-ing/wechat-spider An open-source WeChat crawler that scrapes articles, reading counts, likes, and comments from WeChat official accounts using a man-in-the-m… | 70 | 2942 | active |
| oxylabs/perplexity-scraper A repository of code examples and documentation for Oxylabs' Perplexity Scraper API, which sends prompts to Perplexity and returns AI-gener… | 62 | 2922 | active |
| ChanceYu/front-end-rss An automated RSS aggregator that collects the latest front-end technology articles from popular newsletters and blogs, categorizes them, an… | 77 | 2894 | active |
| scrapinghub/dateparser A Python library that parses human-readable dates in almost any format, including relative expressions like 'two weeks ago', timestamps, an… | 94 | 2853 | stable |
| zzzprojects/html-agility-pack Html Agility Pack (HAP) is a free, open-source HTML parser written in C# that builds a read/write DOM and supports XPath and XSLT queries. … | 95 | 2846 | active |
| spatie/crawler A PHP library by Spatie for crawling links on websites, built on Guzzle promises for concurrent requests. It can execute JavaScript via Chr… | 97 | 2829 | active |
| cv-cat/DouYin_Spider A Python-based Douyin (Chinese TikTok) reverse-engineering toolkit that exposes the platform's full API surface for data collection, live-s… | 62 | 2812 | active |
| ssssssss-team/spider-flow Spider-Flow is a self-hosted Java-based web crawler platform that lets users define scraping workflows visually as flowcharts without writi… | 23 | 11352 | maintenance |
| geziyor/geziyor Geziyor is a fast web crawling and scraping framework for Go, supporting JavaScript rendering via Chrome, caching, proxy management, and au… | 73 | 2775 | active |
| xnl-h4ck3r/waymore waymore is a Python CLI tool that retrieves URLs from multiple web archive and intelligence sources (Wayback Machine, Common Crawl, Alien V… | 90 | 2732 | active |
| microlinkhq/metascraper Metascraper is a Node.js library that extracts unified metadata from any URL by combining Open Graph, JSON-LD, Microdata, RDFa, Twitter Car… | 95 | 2731 | active |
| vladkens/twscrape twscrape is an async Python library and CLI for scraping X/Twitter via its Search and GraphQL endpoints using a pool of your own accounts. … | 96 | 2708 | active |
| Nandaka/PixivUtil2 A Python command-line tool for bulk downloading images from Pixiv and Pixiv FANBOX, with support for downloading by member, tag, bookmark, … | 71 | 2698 | active |
| jae-jae/QueryList QueryList is a progressive PHP web scraping framework built on phpQuery that provides jQuery-like CSS3 DOM selectors and manipulation APIs … | 74 | 2690 | active |
| CharlesPikachu/videodl A lightweight video downloader written in pure Python that parses and downloads videos from dozens of streaming platforms (Douyin, Bilibili… | 83 | 2676 | active |
| chrome-php/chrome A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa… | 89 | 2675 | active |
| cocoindex-io/cocoindex-code A lightweight AST-based semantic code search CLI built on the CocoIndex Rust data transformation engine, with tree-sitter parsing and embed… | 81 | 2675 | active |
| spider-rs/spider Spider is a concurrency-first web crawler and scraper written in Rust that streams pages as they arrive, renders JavaScript only when neede… | 87 | 2672 | active |
| Johnserf-Seed/f2 F2 is an asynchronous Python library and CLI tool for downloading videos and fetching API data from multiple platforms including Douyin, Ti… | 56 | 2620 | active |
| spatie/laravel-sitemap A Laravel package by Spatie that generates XML sitemaps, either by crawling an entire site automatically or by adding URLs manually (includ… | 95 | 2617 | stable |
| brightdata/brightdata-mcp A Model Context Protocol (MCP) server by Bright Data that gives AI agents and LLMs real-time access to public web data through 69 tools cov… | 84 | 2610 | active |
| Serene-Arc/bulk-downloader-for-reddit A Python command-line tool (bdfr) that bulk-downloads and archives Reddit submissions and their media from subreddits, multireddits, users,… | 57 | 2606 | active |
| botswin/BotBrowser BotBrowser is a privacy-focused browser core (Chromium-based) that unifies and controls browser fingerprint signals across platforms, integ… | 85 | 2589 | active |
| apify/fingerprint-suite A modular TypeScript toolkit by Apify for generating realistic browser fingerprints and HTTP headers and injecting them into Playwright or … | 98 | 2581 | active |
| sarperavci/CloudflareBypassForScraping A Python library that bypasses Cloudflare's anti-bot verification for web scraping, supporting cookie generation and request mirroring for … | 72 | 2576 | active |
| lncrawl/lightnovel-crawler Lightnovel Crawler is a Python tool that downloads web novels from 300+ supported sources and converts them into e-books such as EPUB, MOBI… | 97 | 2575 | active |
| simonw/shot-scraper shot-scraper is a Python CLI utility built on Playwright for taking automated screenshots of websites, recording video demos, and scraping … | 89 | 2553 | active |
| guyueyingmu/avbook A self-hosted PHP/Laravel web application that manages a Japanese adult video (JAV) library, backed by crawlers for sites like avmoo, javbu… | 23 | 10036 | maintenance |
| fhamborg/news-please news-please is an open-source Python news crawler and information extractor that pulls structured article data (headline, lead, main text, … | 67 | 2482 | active |
| rust-scraper/scraper A Rust library for parsing HTML documents and querying them with CSS selectors, built on Servo's html5ever and selectors crates for browser… | 89 | 2417 | active |
| scrapinghub/portia Portia is a visual web scraping tool from Scrapinghub built on Scrapy that lets users annotate web pages in a browser to define data extrac… | 10 | 9504 | maintenance |
| OpenBullet OpenBullet 2 is a cross-platform automation suite built on .NET for performing HTTP requests against target web applications and processing… | 85 | 2389 | active |
| gawel/pyquery pyquery is a Python library that provides a jQuery-like API for querying and manipulating XML and HTML documents, built on top of lxml for … | 75 | 2377 | stable |
| apify/agent-skills A collection of production-grade agent skills from Apify that give AI coding agents (Claude Code, Cursor, Windsurf, Codex, Gemini CLI) expe… | 59 | 2361 | active |
| FriendsOfPHP/Goutte Goutte is a PHP screen scraping and web crawling library providing a simple API to crawl websites and extract data from HTML/XML responses.… | 10 | 9192 | maintenance |
| dataabc/weibo-search A Python/Scrapy-based crawler that continuously fetches Weibo keyword and hashtag search results, including full post metadata, images, and… | 71 | 2317 | active |
| sjdirect/abot Abot is an open source C# web crawler framework built for speed and flexibility, handling multithreading, HTTP requests, scheduling, and li… | 74 | 2310 | active |
| 0xMassi/webclaw webclaw is a Rust-based web extraction toolkit that turns any URL into clean, LLM-ready markdown, JSON, or token-optimized text, including … | 77 | 2305 | active |
| Anakin-Inc/anakin AnakinScraper OSS is a self-hosted web scraping API written in Go that turns any website into LLM-ready markdown or structured JSON via a s… | 73 | 2302 | active |
| anaskhan96/soup soup is a small Go library for web scraping with an API modeled after Python's BeautifulSoup. It fetches HTML over HTTP and builds a DOM th… | 96 | 2286 | active |
| saifyxpro/HeadlessX HeadlessX is a self-hosted browser automation and web scraping platform powered by Camoufox (a C++-patched Firefox) to bypass anti-bot syst… | 76 | 2268 | active |
| goclone-dev/goclone Goclone is a Go CLI utility that downloads entire websites to a local directory, preserving relative link structure so the mirrored site ca… | 62 | 2231 | active |
| wabarc/wayback Wayback is an open-source web archiving tool written in Go that captures and preserves web pages via services like Internet Archive, archiv… | 85 | 2227 | active |
| hhursev/recipe-scrapers A Python library for extracting structured recipe data (title, ingredients, instructions, cooking times, images, nutrients) from cooking we… | 98 | 2216 | active |
| ReaJason/xhs A Python SDK that wraps requests to the Xiaohongshu (Little Red Book) web platform for extracting data. It provides a programmatic client f… | 43 | 2202 | active |
| Imangazaliev/DiDOM DiDOM is a fast and simple PHP library for parsing and manipulating HTML and XML documents. It supports loading from strings, files, or URL… | 52 | 2198 | active |
| Owez/yark Yark is a Python CLI tool for archiving YouTube channels, downloading videos and accumulating metadata over time with change reports. It in… | 66 | 2184 | active |
| ericchiang/pup pup is a command line tool for parsing and filtering HTML using CSS selectors, inspired by jq. It reads HTML from stdin, applies selector-b… | 23 | 8435 | maintenance |
| AAndyProgram/SCrawler SCrawler is a Windows GUI application that downloads photos and videos from user profiles across many social media and content sites, inclu… | 93 | 2154 | active |
| philss/floki Floki is an Elixir HTML parser that lets you search document nodes using CSS selectors. It supports multiple parsing backends (mochiweb_htm… | 86 | 2149 | stable |
| Rongronggg9/RSS-to-Telegram-Bot A self-hosted Telegram bot that delivers RSS/Atom feed updates to Telegram chats with rich-text formatting and media support. It is multi-u… | 67 | 2141 | active |
| zorlan/skycaiji SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a … | 77 | 2089 | active |
| oxylabs/how-to-scrape-google-images A Python-based command-line tool that scrapes Google Images search results, including reverse image search based on a provided image URL. I… | 64 | 2055 | active |
| oxylabs/how-to-scrape-google-flights A Python-based free scraper tool and tutorial for extracting flight data (prices, times, airlines) from Google Flights pages, either direct… | 58 | 2048 | active |
| rubycdp/ferrum Ferrum is a Ruby library providing a clean, high-level API to control Chrome or Chromium via the Chrome DevTools Protocol (CDP), with no Se… | 90 | 2037 | active |
| oxylabs/how-to-scrape-amazon-prices A Python-based example repository and free CLI tool for scraping Amazon product prices, best sellers, search results, and deals from depart… | 64 | 2028 | active |
| elliotgao2/gain Gain is an asynchronous web crawling framework for Python built on asyncio, aiohttp, and lxml/pyquery. Users declare items and parsers decl… | 75 | 2019 | active |
| jonhoo/fantoccini Fantoccini is a Rust library providing a high-level async API for programmatically controlling browsers via the WebDriver protocol. It supp… | 75 | 2014 | active |