Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: crawlers

581 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
IonicaBizau/scrape-it
scrape-it is a Node.js web scraping library with a simple, declarative API for extracting data from HTML pages, built on top of tinyreq and…
934073active
fake-useragent/fake-useragent
A Python library that generates realistic, up-to-date browser user-agent strings from a bundled real-world database. It supports random or …
104049active
MikeChongCan/scylla
Scylla is an intelligent proxy pool service that automatically crawls, validates, and serves proxy IPs via a JSON API and web UI. It integr…
344020active
binux/pyspider
pyspider is a powerful web crawler (spider) system written in Python with a built-in WebUI for script editing, task monitoring, project man…
1016774maintenance
kingname/GeneralNewsExtractor
GNE (GeneralNewsExtractor) is a Python library that extracts article title, author, publish time, body text, and images from news page HTML…
813788active
edoardottt/cariddi
Cariddi is a fast command-line web crawler written in Go that takes a list of domains, crawls URLs, and scans for endpoints, secrets, API k…
843753active
Boris-code/feapder
feapder is a powerful, easy-to-use Python web crawling framework offering four spider types (AirSpider, Spider, TaskSpider, BatchSpider) fo…
743735active
JoeanAmier/TikTokDownloader
DouK-Downloader (formerly TikTokDownloader) is an open-source Python tool for downloading and scraping content from Douyin and TikTok, incl…
7115584maintenance
oxylabs/google-ai-mode-scraper
A code repository of examples for Oxylabs' Google AI Mode Scraper, a commercial API that sends prompts to Google AI Mode and returns parsed…
613581active
codelucas/newspaper
newspaper3k is a Python 3 library for discovering, downloading, and parsing news articles from websites. It extracts full text, authors, pu…
6615144maintenance
elliotgao2/toapi
A Python library that turns any website into a JSON API by declaring fields with CSS/XPath selectors. It fetches and parses pages on demand…
753555active
thomasdondorf/puppeteer-cluster
A Node.js library that manages a pool of Puppeteer-controlled Chromium instances for running browser tasks in parallel. It handles job queu…
603514stable
Gerapy/Gerapy
Gerapy is a distributed crawler management framework built on Scrapy, Scrapyd, Django, and Vue.js. It provides a web dashboard for managing…
633512active
google/robotstxt
Google's production robots.txt parser and matcher, released as a C++ library (C++14 compliant). It implements the Robots Exclusion Protocol…
673471stable
oxylabs/free-proxy-list
A repository promoting Oxylabs' free tier of US datacenter proxies (HTTP/HTTPS/SOCKS5), offering 5 US IPs, 20 concurrent sessions, and 5GB …
543434active
any4ai/AnyCrawl
AnyCrawl is a Node.js/TypeScript web crawler and scraping service that converts websites into LLM-ready markdown/JSON data and extracts str…
853415active
opsdisk/pagodo
pagodo is a Python CLI tool that automates passive Google dork searches by scraping the Google Hacking Database (GHDB) and running those qu…
493387active
white0dew/XiaohongshuSkills
A Python CLI tool and agent Skill that automates Xiaohongshu (RED/RedNote) via Chrome DevTools Protocol, supporting publishing posts with i…
603366active
oxylabs/oxylabs-ai-studio-py
A Python SDK for Oxylabs AI Studio, a commercial API offering AI-powered web scraping, crawling, search, site mapping, and browser automati…
593354active
oxylabs/google-news-scraper
A free Python-based command-line tool from Oxylabs that scrapes Google News articles by topic, exporting headlines, URLs, and publication d…
643329active
postaddictme/instagram-php-scraper
A PHP library that scrapes Instagram via its web version to fetch account information, photos, videos, stories, and comments, with optional…
333327active
oxylabs/amazon-scraper
A Python-based free tool and sample code for Oxylabs' Amazon Scraper API that extracts product, search, offer listing, reviews, Q&A, best s…
713326active
tamnd/kage
kage is a Go CLI tool that clones websites into browsable offline folders by rendering each page in headless Chrome, snapshotting the final…
793322active
psf/requests-html
A Python library that combines HTTP requests with intuitive HTML parsing, offering CSS selector and XPath support plus JavaScript rendering…
2313814maintenance
internetarchive/heritrix3
Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler written in Java. It crawls websites and…
953305active
oxylabs/chatgpt-scraper
Code examples and documentation for Oxylabs' ChatGPT Scraper, a commercial Web Scraper API endpoint that sends prompts to ChatGPT and retur…
623301active
php-curl-class/php-curl-class
PHP Curl Class is a PHP library that wraps PHP's cURL extension in a simple object-oriented API for sending HTTP requests. It supports GET,…
933297active
aapatre/Automatic-Udemy-Course-Enroller-GET-PAID-UDEMY-COURSES-for-FREE
A Python script that scrapes coupon sites for free Udemy course coupons and automatically enrolls the user in paid courses for free using S…
583291active
apache/nutch
Apache Nutch is a highly extensible and scalable open-source web crawler built on Apache Hadoop data structures. It supports batch crawling…
773277stable
itsOwen/CyberScraper-2077
CyberScraper 2077 is an AI-powered web scraping application with a Streamlit GUI that uses OpenAI, Gemini, or local Ollama models to intell…
673245active
jsvine/waybackpack
Waybackpack is a Python command-line tool that downloads the entire Wayback Machine archive for a given URL, saving every archived snapshot…
303228active
oxylabs/ai-crawler-py
A Python client library for Oxylabs AI Studio's AI-Crawler, a commercial service that crawls websites from a starting URL, uses natural lan…
613218active
s0md3v/Photon
Photon is a fast Python-based web crawler designed for OSINT (open-source intelligence) tasks. It extracts URLs, emails, social media accou…
6613146maintenance
devanshbatham/ParamSpider
ParamSpider is a Python CLI tool that mines URLs for a domain or list of domains from web archives (Wayback Machine), filtering out uninter…
553160active
omkarcloud/google-maps-scraper
A desktop application and API for scraping Google Maps business data, extracting 50+ data points including emails, phone numbers, social pr…
723130active
scrapy/scrapyd
Scrapyd is a service daemon for deploying and running Scrapy spiders. It lets you upload Scrapy projects and control spiders via a JSON HTT…
663101stable
symfony/panther
Symfony Panther is a PHP library for browser testing and web scraping that drives real browsers (Chrome, Firefox) via the W3C WebDriver pro…
733067active
Qianlitp/crawlergo
crawlergo is a Go-based browser crawler that uses headless Chrome to discover URLs for web vulnerability scanners. It renders pages, fills …
273034active
adryfish/fingerprint-chromium
A fingerprint browser built on Ungoogled Chromium that lets users spoof or control browser fingerprint characteristics to avoid detection. …
762991active
5ime/video_spider
A PHP-based web service that parses short-video links from platforms like Douyin, Kuaishou, Weibo, and Pipixia to return watermark-free vid…
572975active
nashsu/AutoCLI
AutoCLI is a blazing-fast, memory-safe command-line tool written in Rust that fetches information from 55+ websites (Twitter/X, Reddit, You…
652950active
rust-headless-chrome/rust-headless-chrome
A Rust library providing a high-level API to control headless Chrome or Chromium via the DevTools Protocol, serving as the Rust equivalent …
812947active
striver-ing/wechat-spider
An open-source WeChat crawler that scrapes articles, reading counts, likes, and comments from WeChat official accounts using a man-in-the-m…
702942active
oxylabs/perplexity-scraper
A repository of code examples and documentation for Oxylabs' Perplexity Scraper API, which sends prompts to Perplexity and returns AI-gener…
622922active
ChanceYu/front-end-rss
An automated RSS aggregator that collects the latest front-end technology articles from popular newsletters and blogs, categorizes them, an…
772894active
scrapinghub/dateparser
A Python library that parses human-readable dates in almost any format, including relative expressions like 'two weeks ago', timestamps, an…
942853stable
zzzprojects/html-agility-pack
Html Agility Pack (HAP) is a free, open-source HTML parser written in C# that builds a read/write DOM and supports XPath and XSLT queries. …
952846active
spatie/crawler
A PHP library by Spatie for crawling links on websites, built on Guzzle promises for concurrent requests. It can execute JavaScript via Chr…
972829active
cv-cat/DouYin_Spider
A Python-based Douyin (Chinese TikTok) reverse-engineering toolkit that exposes the platform's full API surface for data collection, live-s…
622812active
ssssssss-team/spider-flow
Spider-Flow is a self-hosted Java-based web crawler platform that lets users define scraping workflows visually as flowcharts without writi…
2311352maintenance
geziyor/geziyor
Geziyor is a fast web crawling and scraping framework for Go, supporting JavaScript rendering via Chrome, caching, proxy management, and au…
732775active
xnl-h4ck3r/waymore
waymore is a Python CLI tool that retrieves URLs from multiple web archive and intelligence sources (Wayback Machine, Common Crawl, Alien V…
902732active
microlinkhq/metascraper
Metascraper is a Node.js library that extracts unified metadata from any URL by combining Open Graph, JSON-LD, Microdata, RDFa, Twitter Car…
952731active
vladkens/twscrape
twscrape is an async Python library and CLI for scraping X/Twitter via its Search and GraphQL endpoints using a pool of your own accounts. …
962708active
Nandaka/PixivUtil2
A Python command-line tool for bulk downloading images from Pixiv and Pixiv FANBOX, with support for downloading by member, tag, bookmark, …
712698active
jae-jae/QueryList
QueryList is a progressive PHP web scraping framework built on phpQuery that provides jQuery-like CSS3 DOM selectors and manipulation APIs …
742690active
CharlesPikachu/videodl
A lightweight video downloader written in pure Python that parses and downloads videos from dozens of streaming platforms (Douyin, Bilibili…
832676active
chrome-php/chrome
A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa…
892675active
cocoindex-io/cocoindex-code
A lightweight AST-based semantic code search CLI built on the CocoIndex Rust data transformation engine, with tree-sitter parsing and embed…
812675active
spider-rs/spider
Spider is a concurrency-first web crawler and scraper written in Rust that streams pages as they arrive, renders JavaScript only when neede…
872672active
Johnserf-Seed/f2
F2 is an asynchronous Python library and CLI tool for downloading videos and fetching API data from multiple platforms including Douyin, Ti…
562620active
spatie/laravel-sitemap
A Laravel package by Spatie that generates XML sitemaps, either by crawling an entire site automatically or by adding URLs manually (includ…
952617stable
brightdata/brightdata-mcp
A Model Context Protocol (MCP) server by Bright Data that gives AI agents and LLMs real-time access to public web data through 69 tools cov…
842610active
Serene-Arc/bulk-downloader-for-reddit
A Python command-line tool (bdfr) that bulk-downloads and archives Reddit submissions and their media from subreddits, multireddits, users,…
572606active
botswin/BotBrowser
BotBrowser is a privacy-focused browser core (Chromium-based) that unifies and controls browser fingerprint signals across platforms, integ…
852589active
apify/fingerprint-suite
A modular TypeScript toolkit by Apify for generating realistic browser fingerprints and HTTP headers and injecting them into Playwright or …
982581active
sarperavci/CloudflareBypassForScraping
A Python library that bypasses Cloudflare's anti-bot verification for web scraping, supporting cookie generation and request mirroring for …
722576active
lncrawl/lightnovel-crawler
Lightnovel Crawler is a Python tool that downloads web novels from 300+ supported sources and converts them into e-books such as EPUB, MOBI…
972575active
simonw/shot-scraper
shot-scraper is a Python CLI utility built on Playwright for taking automated screenshots of websites, recording video demos, and scraping …
892553active
guyueyingmu/avbook
A self-hosted PHP/Laravel web application that manages a Japanese adult video (JAV) library, backed by crawlers for sites like avmoo, javbu…
2310036maintenance
fhamborg/news-please
news-please is an open-source Python news crawler and information extractor that pulls structured article data (headline, lead, main text, …
672482active
rust-scraper/scraper
A Rust library for parsing HTML documents and querying them with CSS selectors, built on Servo's html5ever and selectors crates for browser…
892417active
scrapinghub/portia
Portia is a visual web scraping tool from Scrapinghub built on Scrapy that lets users annotate web pages in a browser to define data extrac…
109504maintenance
OpenBullet
OpenBullet 2 is a cross-platform automation suite built on .NET for performing HTTP requests against target web applications and processing…
852389active
gawel/pyquery
pyquery is a Python library that provides a jQuery-like API for querying and manipulating XML and HTML documents, built on top of lxml for …
752377stable
apify/agent-skills
A collection of production-grade agent skills from Apify that give AI coding agents (Claude Code, Cursor, Windsurf, Codex, Gemini CLI) expe…
592361active
FriendsOfPHP/Goutte
Goutte is a PHP screen scraping and web crawling library providing a simple API to crawl websites and extract data from HTML/XML responses.…
109192maintenance
dataabc/weibo-search
A Python/Scrapy-based crawler that continuously fetches Weibo keyword and hashtag search results, including full post metadata, images, and…
712317active
sjdirect/abot
Abot is an open source C# web crawler framework built for speed and flexibility, handling multithreading, HTTP requests, scheduling, and li…
742310active
0xMassi/webclaw
webclaw is a Rust-based web extraction toolkit that turns any URL into clean, LLM-ready markdown, JSON, or token-optimized text, including …
772305active
Anakin-Inc/anakin
AnakinScraper OSS is a self-hosted web scraping API written in Go that turns any website into LLM-ready markdown or structured JSON via a s…
732302active
anaskhan96/soup
soup is a small Go library for web scraping with an API modeled after Python's BeautifulSoup. It fetches HTML over HTTP and builds a DOM th…
962286active
saifyxpro/HeadlessX
HeadlessX is a self-hosted browser automation and web scraping platform powered by Camoufox (a C++-patched Firefox) to bypass anti-bot syst…
762268active
goclone-dev/goclone
Goclone is a Go CLI utility that downloads entire websites to a local directory, preserving relative link structure so the mirrored site ca…
622231active
wabarc/wayback
Wayback is an open-source web archiving tool written in Go that captures and preserves web pages via services like Internet Archive, archiv…
852227active
hhursev/recipe-scrapers
A Python library for extracting structured recipe data (title, ingredients, instructions, cooking times, images, nutrients) from cooking we…
982216active
ReaJason/xhs
A Python SDK that wraps requests to the Xiaohongshu (Little Red Book) web platform for extracting data. It provides a programmatic client f…
432202active
Imangazaliev/DiDOM
DiDOM is a fast and simple PHP library for parsing and manipulating HTML and XML documents. It supports loading from strings, files, or URL…
522198active
Owez/yark
Yark is a Python CLI tool for archiving YouTube channels, downloading videos and accumulating metadata over time with change reports. It in…
662184active
ericchiang/pup
pup is a command line tool for parsing and filtering HTML using CSS selectors, inspired by jq. It reads HTML from stdin, applies selector-b…
238435maintenance
AAndyProgram/SCrawler
SCrawler is a Windows GUI application that downloads photos and videos from user profiles across many social media and content sites, inclu…
932154active
philss/floki
Floki is an Elixir HTML parser that lets you search document nodes using CSS selectors. It supports multiple parsing backends (mochiweb_htm…
862149stable
Rongronggg9/RSS-to-Telegram-Bot
A self-hosted Telegram bot that delivers RSS/Atom feed updates to Telegram chats with rich-text formatting and media support. It is multi-u…
672141active
zorlan/skycaiji
SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a …
772089active
oxylabs/how-to-scrape-google-images
A Python-based command-line tool that scrapes Google Images search results, including reverse image search based on a provided image URL. I…
642055active
oxylabs/how-to-scrape-google-flights
A Python-based free scraper tool and tutorial for extracting flight data (prices, times, airlines) from Google Flights pages, either direct…
582048active
rubycdp/ferrum
Ferrum is a Ruby library providing a clean, high-level API to control Chrome or Chromium via the Chrome DevTools Protocol (CDP), with no Se…
902037active
oxylabs/how-to-scrape-amazon-prices
A Python-based example repository and free CLI tool for scraping Amazon product prices, best sellers, search results, and deals from depart…
642028active
elliotgao2/gain
Gain is an asynchronous web crawling framework for Python built on asyncio, aiohttp, and lxml/pyquery. Users declare items and parsers decl…
752019active
jonhoo/fantoccini
Fantoccini is a Rust library providing a high-level async API for programmatically controlling browsers via the WebDriver protocol. It supp…
752014active

← prev page 2 / 6 next →