Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: crawlers

581 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
watercrawl/WaterCrawl
WaterCrawl is a self-hostable web application (Python/Django/Scrapy/Celery) that crawls websites and transforms web content into LLM-ready …
822010active
supermemoryai/markdowner
Markdowner is a fast web service that converts any website into LLM-ready markdown, with optional LLM filtering, detailed responses, and au…
241999active
nottelabs/notte
Notte is a full-stack framework and cloud platform for building, deploying, and scaling AI web agents and browser automations. It combines …
841997active
zhegexiaohuozi/SeimiCrawler
SeimiCrawler is an agile, standalone, distributed Java crawler framework inspired by Python's Scrapy, with deep Spring Boot integration and…
851990active
AWeirdDev/flights
fast-flights is a Python library that scrapes Google Flights by generating Base64-encoded Protobuf query strings, returning strongly-typed …
861939active
feder-cr/invisible_playwright
A Python library that provides an antidetect, stealth-patched Firefox build for Playwright, with fingerprints set at the C++ engine level a…
811938active
damoeb/rss-proxy
RSS-proxy is a self-hostable web service that generates RSS, ATOM, or JSON feeds from almost any static website by analyzing its HTML struc…
321924active
MarginaliaSearch/MarginaliaSearch
Marginalia Search is an independent, open-source internet search engine that indexes text-oriented, non-commercial, small and old websites.…
651923active
KEV0143/Parser-Chitai-Gorod
A Python-based scraper for the Russian online bookstore Chitai-Gorod that collects book URLs across catalog pages and extracts structured p…
291914active
extractus/article-extractor
A TypeScript library that extracts the main article content, title, image, and metadata from a given URL or raw HTML string. It supports cu…
981909active
trevorhobenshield/twitter-api-client
A Python library implementing X/Twitter's v1, v2, and GraphQL APIs for automation and scraping. It supports account actions like tweeting, …
301894active
404-novel-project/novel-downloader
An extensible userscript (Tampermonkey/Greasemonkey/Violentmonkey) that downloads novels from many Chinese web novel sites and exports them…
761877active
coder-hxl/x-crawl
x-crawl is a flexible Node.js crawler library that supports crawling dynamic pages, static pages, API data, and files, with optional AI ass…
661877active
ThePhaseless/Byparr
Byparr is a self-hosted Python service that solves antibot browser challenges (like Cloudflare checks) and returns valid clearance cookies …
901865active
sarperavci/GoogleRecaptchaBypass
A Python library that automatically solves Google reCAPTCHA v2 challenges in under five seconds using browser automation with DrissionPage …
681855active
microlinkhq/browserless
A Node.js library that wraps Puppeteer to provide a production-ready headless Chrome/Chromium driver with built-in screenshot, PDF generati…
951831active
tryolabs/requestium
Requestium is a Python library that merges Requests, Selenium, and Parsel into a single integrated tool for web automation. It lets scripts…
771830active
enetx/surf
Surf is an advanced HTTP client library for Go with fluent, chainable API design. It supports browser impersonation (Chrome/Firefox), JA3/J…
831808active
deweizhu/bookget
bookget is a Go-based command-line tool for downloading digitized ancient books and rare texts from 50+ digital libraries. It ships prebuil…
601764active
LoseNine/ruyipage
RuyiPage is a Python browser automation framework built on Firefox and the WebDriver BiDi protocol, shipping with an anti-detection Firefox…
791759active
josh0xA/darkdump
Darkdump is an open-source OSINT tool for querying multiple dark web search engines and scraping onion site results for emails, metadata, k…
701757active
website-scraper/node-website-scraper
A Node.js library that downloads entire websites to a local directory, including HTML, CSS, images, and JavaScript assets. It parses HTTP r…
721751active
egoist/sitefetch
A Node.js CLI tool that crawls an entire website and saves its pages as a single text file, using Mozilla Readability to extract clean cont…
221736active
claffin/cloudproxy
CloudProxy is a self-hosted Python tool that provisions and manages proxy servers across multiple cloud providers, rotating IPs to improve …
811722active
Jules-WinnfieldX/CyberDropDownloader
A Python-based bulk downloader that scrapes and downloads files from Cyberdrop.me and dozens of other file hosts and image galleries. It is…
101720active
3441293738/creatorhub
CreatorHub is a self-hosted web panel built with Python and FastAPI for managing, monitoring, scraping, downloading, and publishing content…
581712active
Python3Spiders/WeiboSuperSpider
A Weibo (Chinese microblog) scraping toolbox in Python covering users, topics, and comments, with extras like image downloading, sentiment …
751705active
oxylabs/google-play-scraper
A free Python-based Google Play Store scraper that collects public app, movie, and book data via search queries. It is a companion tool to …
671693active
vibheksoni/stealth-browser-mcp
A Python MCP server that exposes stealth browser automation (via nodriver and Chrome DevTools Protocol) to AI agents, letting them navigate…
611674active
rushter/selectolax
Selectolax is a fast Python HTML5 parser library written in Cython, binding to the Modest and Lexbor parsing engines. It provides CSS selec…
941665active
MgArcher/Text_select_captcha
A PyTorch-based deep learning system that recognizes click-based (text-select) CAPTCHAs by detecting and ordering Chinese character positio…
691656active
srx-2000/spider_collection
A collection of Python web crawler scripts targeting sites like Bilibili, Zhihu, Weibo, NetEase Music, GitHub, and Anjuke, built with reque…
321646active
Jesseovo/last30days-skill-cn
An AI Agent skill (for Claude Code / OpenClaw) that automatically searches content from the last 30 days across 8 major Chinese internet pl…
561642active
wu529778790/panhub.shenzjd.com
PanHub is a self-hostable netdisk search aggregator that combines results from Quark, Aliyun Drive, Baidu Netdisk, 115, Thunder and 80+ Tel…
621617active
hartator/wayback-machine-downloader
A Ruby command-line tool that downloads an entire website from the Internet Archive Wayback Machine, restoring original files and directory…
235930maintenance
ArchiveTeam/grab-site
grab-site is a preconfigured web crawler for archiving websites, producing WARC files via a fork of wpull. It includes a dashboard for moni…
431607active
nerevu/riko
riko is a pure Python stream processing library modeled after Yahoo! Pipes, combining reusable, configuration-driven modular pipes with syn…
881606active
matthewmueller/x-ray
x-ray is a Node.js web scraping library that lets you define flexible schemas to structure data from any website using jQuery-like selector…
665907maintenance
Altimis/Scweet
Scweet is a Python library and CLI for scraping tweets, profile timelines, followers, following lists, and user profiles from Twitter/X wit…
781604active
lanyeeee/jmcomic-downloader
A multi-threaded GUI downloader for the 18comic.vip (jmcomic) manga site, built with Tauri (Rust backend, Vue frontend). It supports search…
891587active
s0md3v/uro
uro is a Python CLI tool that declutters URL lists for crawling and security testing without making any HTTP requests. It removes duplicate…
291587stable
xnl-h4ck3r/xnLinkFinder
xnLinkFinder is a Python CLI tool that discovers endpoints, potential parameters, target-specific wordlists, and secrets for a given target…
781585active
ttttmr/Wechat2RSS
Wechat2RSS is a service and self-hostable tool that converts WeChat official account (公众号) articles into RSS feeds, aiming for updates with…
751557active
ulixee/hero
Hero is a headless web browser built specifically for web scraping, powered by Chrome and controlled from NodeJS with a fully compliant DOM…
691554active
Rhizome-Conifer/conifer
Conifer is an open-source web archiving platform for capturing, replaying, and sharing collections of archived web pages through a user-fri…
741549active
skernelx/tavily-key-generator
A Python toolkit that automates signup flows for Tavily, Firecrawl, and Exa using real browser automation (Playwright/Camoufox), Turnstile …
481549active
yujiosaka/headless-chrome-crawler
A Node.js library providing a distributed web crawler powered by Headless Chrome via Puppeteer. It can crawl JavaScript-rendered (SPA) webs…
235635maintenance
oxylabs/ai-map-py
AI-Map is a Python SDK client for Oxylabs AI Studio's AI-powered website mapping service, which discovers and extracts relevant URLs from a…
501537active
LifeActor/ykdl
YouKuDownLoader (ykdl) is a Python command-line video downloader focused on China mainland video sites, forked from you-get with restructur…
471533active
justfoolingaround/animdl
animdl is a lightweight Python CLI tool that scrapes, streams, and downloads anime episodes from supported providers. It supports quality s…
321522active
tidyverse/rvest
rvest is an R package from the tidyverse for scraping (harvesting) data from web pages, inspired by Beautiful Soup and RoboBrowser. It prov…
431520active
Danny-Dasilva/CycleTLS
CycleTLS is a Go library with a JavaScript/TypeScript wrapper that lets clients spoof TLS/JA3 (and JA4) fingerprints when making HTTP reque…
731518active
SpiderClub/haipproxy
A high-availability distributed IP proxy pool built with Scrapy and Redis that scrapes free proxies from the internet, validates them, and …
235523maintenance
oxylabs/browser-agent-py
A Python SDK for Oxylabs AI Studio's Browser Agent, a cloud service that automates real-user browsing tasks (clicking, typing, scrolling, s…
511505active
requests-cache/requests-cache
requests-cache is a persistent HTTP caching library for Python's requests library, providing a drop-in CachedSession and optional global pa…
771501stable
JustAnotherArchivist/snscrape
snscrape is a Python-based scraper for social networking services that extracts posts, profiles, hashtags, and search results from platform…
325444maintenance
roach-php/core
Roach is a complete web scraping and crawling toolkit for PHP, heavily inspired by Python's Scrapy. It lets developers define spiders that …
451455active
submato/xhscrawl
A Python-based reverse-engineering toolkit for Xiaohongshu (XHS) web APIs, focusing on generating the encrypted x-s signature parameter via…
721452active
tinyfish-io/agentql
AgentQL is a suite of tools for extracting structured data and automating workflows on live websites using an AI-powered natural language q…
701451active
scrapy-plugins/scrapy-playwright
A Scrapy download handler that uses Playwright for Python to fetch pages, enabling scraping of JavaScript-rendered sites while keeping the …
861439active
ShilongLee/Crawler
A self-hostable crawler API server that exposes HTTP endpoints for scraping public data from Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo…
531435active
drawrowfly/tiktok-scraper
A TypeScript library and CLI tool that scrapes TikTok metadata from user, hashtag, trend, and music pages and downloads video posts without…
235173maintenance
zhaoolee/garss
Garss (嘎!RSS) is a self-hosted RSS aggregation and reading system that uses GitHub Actions to collect hundreds of RSS feeds and render them…
821426active
rebrowser/rebrowser-patches
A collection of source-code patches for Puppeteer and Playwright that fix automation leaks and help avoid bot detection systems like Cloudf…
341424active
cdpdriver/zendriver
Zendriver is an async-first Python web scraping and browser automation framework built on the Chrome Devtools Protocol, forked from nodrive…
881409active
KoalaBear84/OpenDirectoryDownloader
A cross-platform C#/.NET command-line tool that indexes open directory listings across 130+ supported formats, including FTP(S), Google Dri…
941389active
lorey/mlscraper
mlscraper is a Python library that automatically extracts structured data from HTML pages using machine learning. Instead of writing CSS se…
231385active
Avnsx/fansly-downloader
A Python-based tool for bulk downloading photos, videos, and audio from fansly.com, also shipped as a standalone Windows executable. It sup…
101384active
CIRCL/AIL-framework
AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T…
671378active
mattsse/chromiumoxide
chromiumoxide is a Rust library providing a high-level async API for controlling Chrome or Chromium via the Chrome DevTools Protocol. It ca…
751375active
okfn-brasil/querido-diario
Querido Diário is an open-source project by Open Knowledge Brasil that scrapes and aggregates Brazilian municipal official gazettes (diário…
751373active
xiaohucode/xiangse
A curated collection of video and manga source plugins (.xbs files) for the Xiangse Guige (香色闺阁) reading/media app, imported via URL. It ag…
101359active
scrapy/parsel
Parsel is a BSD-licensed Python library for extracting data from HTML, XML, and JSON documents using CSS selectors, XPath expressions, JMES…
881352active
raznem/parsera
Parsera is a lightweight Python library for scraping websites using LLMs, letting users define elements to extract with natural-language de…
541350active
LeetaoGoooo/RSSAid
RSSAid is a Flutter-based mobile app that complements RSSHub by helping users discover and subscribe to RSS feeds from websites, similar to…
761345active
philippta/flyscrape
Flyscrape is a standalone command-line web scraping tool written in Go that lets users write extraction logic in JavaScript with a jQuery-l…
421345active
SpiderClub/weibospider
A distributed web crawler for Sina Weibo (Chinese microblogging platform) built with Python, Celery, and requests. It scrapes user profiles…
324793maintenance
mvdbos/php-spider
A configurable and extensible PHP web spider library for crawling websites. It supports breadth-first and depth-first traversal, URI discov…
751341active
tophubs/TopList
TopList (今日热榜) is a self-hosted aggregation website that collects trending headlines from popular sites like Zhihu, Hupu, and V2EX. It is w…
324730maintenance
karust/openserp
OpenSERP is a self-hosted, MIT-licensed SERP API and CLI written in Go that returns structured search results from Google, Bing, Yandex, Ba…
911306active
yasserg/crawler4j
crawler4j is an open-source web crawler library for Java that provides a simple interface for building multi-threaded web crawlers in minut…
234618maintenance
dwisiswant0/go-dork
go-dork is a fast command-line dork scanner written in Go that automates Google dorking across multiple search engines. It supports Google,…
231301stable
bookstairs/bookhunter
bookhunter is a Go command-line tool for scraping and downloading ebooks from sources like Talebook, SoBooks, Telegram channels, and China'…
541295active
techwithtim/Price-Tracking-Web-Scraper
A full-stack price tracking application that scrapes product prices (currently Amazon.ca) using Playwright and Bright Data's Scraping Brows…
291294active
zohaibbashir/Google-Maps-Scrapper
A Python CLI script built on Playwright that scrapes Google Maps listings to extract business details such as name, address, website, phone…
651289active
oxylabs/paid-proxy-servers
A promotional GitHub repository for Oxylabs' commercial paid proxy services, covering residential, mobile, datacenter, ISP, and SOCKS5 prox…
591289active
TheBeastLT/torrentio-scraper
Torrentio is a Stremio addon ecosystem that scrapes public torrent providers and serves the results as Stremio stream results. The reposito…
771280active
sardanioss/httpcloak
httpcloak is a Go HTTP client library that reproduces browser-identical TLS, HTTP/2, and HTTP/3 fingerprints (JA3/JA4, Akamai, header order…
611276active
kkangert/kspider
Kspider is a self-hosted visual web scraping platform written in Java where users define crawler workflows as flowcharts without writing ba…
141269active
minsight-ai-info/AI-Search-Hub
AI Search Hub is an open-source Skill that aggregates native AI search capabilities from platforms like Gemini, Grok, Doubao, and Yuanbao i…
501250active
egbertbouman/youtube-comment-downloader
A Python script and library for downloading YouTube video comments without using the official YouTube API. It outputs comments in JSONL, JS…
881247active
Decodo/Decodo
Decodo (formerly Smartproxy) is a commercial rotating proxy network and web scraping platform offering 125M+ residential, mobile, ISP, and …
691238active
raawaa/jav-scrapy
A TypeScript-based Node.js CLI tool that batch-scrapes JAV (adult video) metadata, magnet links, and cover images from source websites. It …
941235active
eatmoreduck/boss-zhipin-scraper
A Python CLI scraper for BOSS Zhipin (zhipin.com) that connects to a locally logged-in Chrome via the Chrome DevTools Protocol to call the …
571231active
iszhouhua/social-media-copilot
An open-source browser extension (built with WXT and TypeScript) that scrapes data from Chinese social media platforms including Xiaohongsh…
651228active
daijro/browserforge
BrowserForge is a Python library that generates realistic browser headers and fingerprints, mimicking real-world browser, OS, and device di…
561224active
firecrawl/web-agent
An open-source TypeScript framework for building autonomous web research agents, layered from a Next.js chat template down to an agent core…
501222active
cnbattle/douyin
A Go-based crawler that scrapes Douyin (TikTok China) recommendation and search page video lists by controlling the mobile app on a real de…
611219active
Silent1566/OmniBox-Spider
A collection of spider (scraper) sources and interfaces for the OmniBox media application, aggregated from publicly available internet info…
581213active
lexmount/moli
Moli is a lightweight, fast headless browser built in Rust (on Servo technology) designed for AI agents to fetch, render, and extract web p…
791206active

← prev page 3 / 6 next →