Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: crawlers

581 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
Firecrawl
Firecrawl is an open-source web scraping and crawling API that turns websites into clean Markdown, structured JSON, screenshots, and other …
88172809active
unclecode/crawl4ai
Crawl4AI is an open-source Python library that crawls websites with a headless browser and converts pages into clean, LLM-ready Markdown fo…
8979471active
D4Vinci/Scrapling
Scrapling is an adaptive Python web scraping framework that handles everything from single requests to full-scale concurrent crawls. It fea…
8976660active
Panniantong/Agent-Reach
Agent Reach is a Python CLI and MCP tool that gives AI agents the ability to read and search across platforms like Twitter, Reddit, YouTube…
7875643active
scrapy/scrapy
Scrapy is a fast, high-level web crawling and scraping framework for Python used to extract structured data from websites. It provides a fu…
9964048stable
NanmiCoder/MediaCrawler
MediaCrawler is a Python-based multi-platform social media crawler that scrapes notes, videos, posts, and comments from Xiaohongshu, Douyin…
7363826active
soimort/you-get
You-Get is a tiny command-line utility written in Python to download media content (videos, audios, images) from the web when no other hand…
6756871active
NaiboWang/EasySpider
EasySpider is a free, open-source visual no-code web crawler and browser automation (RPA) tool where users design scraping tasks by clickin…
8744442active
iawia002/lux
Lux is a fast and simple video downloader written in Go, available both as a CLI tool and a library. It supports downloading videos from ma…
5631656active
CloakHQ/CloakBrowser
CloakBrowser is a stealth Chromium distribution with 73 source-level C++ fingerprint patches that passes major bot-detection systems like C…
8130849active
feder-cr/Jobs_Applier_AI_Agent_AIHawk
AIHawk is an open-source Python AI agent that automates job applications by scraping job postings via browser automation and auto-applying …
5830260active
ScrapeGraphAI/Scrapegraph-ai
ScrapeGraphAI is a Python library that uses LLMs and direct graph logic to build web scraping pipelines for websites and local documents (X…
8829959active
ArchiveBox/ArchiveBox
ArchiveBox is an open-source, self-hosted web archiving application that saves URLs, browser history, bookmarks, RSS feeds, and social medi…
8328188active
Crawlee
Crawlee is a web scraping and browser automation library for Node.js/TypeScript (with a Python port) for building reliable crawlers. It int…
9825516stable
gocolly/colly
Colly is a fast and elegant web scraping and crawling framework for Go. It provides a clean callback-based API for making HTTP requests, pa…
6625483stable
jhao104/proxy_pool
A Python proxy pool service that periodically scrapes free proxies from 15+ sources, validates them, and stores them in Redis/SSDB. It expo…
6223642active
BuilderIO/gpt-crawler
A Node.js/TypeScript crawler that scrapes one or more websites and generates knowledge files (JSON) for creating custom GPTs or OpenAI assi…
3122393active
h4ckf0r0day/obscura
Obscura is an open-source headless browser engine written in Rust that runs JavaScript via V8 and speaks the Chrome DevTools Protocol, acti…
8122334active
xifangczy/cat-catch
Cat-catch is an open-source browser extension that sniffs and lists media resources (video, audio, images) loaded by the current web page. …
9721548active
Evil0ctal/Douyin_TikTok_Download_API
A high-performance asynchronous Python service that scrapes and parses data from Douyin, TikTok, Kuaishou, and Bilibili, exposing a FastAPI…
5519679active
mikf/gallery-dl
gallery-dl is a command-line program to download image galleries and collections from many image hosting sites such as Pixiv, Danbooru, Dev…
9619332active
projectdiscovery/katana
Katana is a fast, configurable web crawling and spidering framework written in Go, supporting both standard HTTP-based and headless browser…
9417348active
getmaxun/maxun
Maxun is an open-source no-code platform for web scraping, crawling, search, and AI-powered data extraction that turns websites into struct…
9417299active
MODSetter/SurfSense
SurfSense is an open-source, self-hostable NotebookLM alternative that combines a personal knowledge base with live open-web research conne…
9016017active
FlareSolverr/FlareSolverr
FlareSolverr is a proxy server that bypasses Cloudflare and DDoS-GUARD protection by solving browser challenges with Selenium and undetecte…
9515316active
PuerkitoBio/goquery
goquery is a Go library that provides a jQuery-like syntax for parsing, traversing, and manipulating HTML documents, built on Go's net/html…
8614981stable
browserless/browserless
Browserless is a platform for deploying and managing headless browsers (Chrome, Firefox, WebKit) in Docker, usable self-hosted or via their…
9913634active
pystardust/ani-cli
ani-cli is a POSIX shell script CLI tool for browsing and watching anime from the terminal, scraping the anidb site and playing videos via …
9713602active
instaloader/instaloader
Instaloader is a Python command-line tool and library for downloading pictures, videos, captions, comments, and metadata from Instagram. It…
8713240active
ultrafunkamsterdam/undetected-chromedriver
A Python library that patches Selenium's ChromeDriver binary so automated Chrome sessions avoid detection by anti-bot systems like Cloudfla…
4612808active
JoeanAmier/XHS-Downloader
A tool for extracting links and downloading content (images, videos, live photos) from XiaoHongShu (RedNote) posts. It ships as a Python TU…
7612489active
g1879/DrissionPage
DrissionPage is a Python-based web automation library that combines browser control (via Chromium's DevTools Protocol, without WebDriver) w…
7012405active
crawlab-team/crawlab
Crawlab is a Go-based distributed web crawler management platform with a web UI for managing, scheduling, and monitoring spiders written in…
5212262active
jina-ai/reader
Jina AI Reader converts any URL into LLM-friendly markdown via the r.jina.ai prefix, and searches the web into markdown via s.jina.ai. It r…
6211912active
code4craft/webmagic
WebMagic is a scalable web crawler framework for Java covering the full crawl lifecycle: downloading, URL management, content extraction (X…
6011678active
daijro/camoufox
Camoufox is an open-source anti-detect browser built on Firefox for web scraping and AI agents, with browser fingerprint spoofing and Playw…
9011455active
jhy/jsoup
jsoup is a Java library for parsing, manipulating, and cleaning real-world HTML and XML, implementing the WHATWG HTML5 specification to pro…
9511387stable
ihmily/DouyinLiveRecorder
A Python-based livestream recording application that can run unattended in a loop and record multiple streams simultaneously from 40+ platf…
6410786active
pinchtab/pinchtab
PinchTab is a standalone Go HTTP server that gives AI agents direct control over Chrome via CDP, with stealth injection and multi-instance …
8210144active
dataabc/weiboSpider
A Python command-line crawler that scrapes posts and profile data from Sina Weibo users, writing results to txt/csv/json files or MySQL/Mon…
629695active
RSS-Bridge/rss-bridge
RSS-Bridge is a PHP web application that generates RSS, Atom, and JSON feeds for websites that don't offer one, using hundreds of site-spec…
759190active
kepano/defuddle
Defuddle is a TypeScript library, CLI, and hosted service that extracts the main content from web pages, removing clutter like comments, si…
879166active
qiye45/wechatDownload
A desktop tool for batch downloading WeChat Official Account (公众号) articles, saving them as html/mhtml/md/pdf/docx/csv and preserving embed…
909105active
Thysrael/Horizon
Horizon is an AI-powered news aggregation platform that fetches content from sources like Hacker News, GitHub, RSS, Reddit, and Telegram, s…
599040active
HerbertHe/iptv-sources
A service that automatically aggregates and updates IPTV m3u playlist sources from multiple public repositories, with EPG data and optional…
718921active
jo-inc/camofox-browser
A self-hosted anti-detection browser server for AI agents, wrapping the Camoufox Firefox fork that spoofs fingerprints at the C++ level. It…
828894active
eze-is/web-access
An Agent Skill that gives AI coding agents (Claude Code, Cursor, Gemini CLI, etc.) full web access capabilities: three-tier channel dispatc…
658745active
hardikvasa/google-images-download
A Python command-line tool that searches and downloads hundreds of images from Google Images to local storage. It uses Selenium with Chrome…
708684active
kangvcar/InfoSpider
InfoSpider is an open-source Python toolbox that crawls a user's own personal data from dozens of Chinese and international services (email…
588247active
alirezamika/autoscraper
AutoScraper is a Python library that automatically learns scraping rules from a URL or HTML content plus a list of sample data you want to …
667904stable
andeya/pholcus
Pholcus is a distributed, high-concurrency web crawler framework written in pure Go. It supports standalone, server, and client modes with …
907577active
mgdm/htmlq
htmlq is a command-line tool, like jq but for HTML, that extracts content from HTML documents using CSS selectors. Written in Rust, it read…
617576stable
Steel Browser
Steel Browser is an open-source browser API and sandbox that manages browser sessions, proxies, stealth, and lifecycle so developers can bu…
897545active
cv-cat/Spider_XHS
A Python library that reverse-engineers Xiaohongshu (Little Red Book) signature algorithms and wraps the platform's PC, creator, and Pugong…
887423active
berstend/puppeteer-extra
A modular plugin framework for Puppeteer (and Playwright via playwright-extra) that extends headless browser automation with drop-in plugin…
327398active
autoscrape-labs/pydoll
Pydoll is a Python library for automating Chromium-based browsers directly over the Chrome DevTools Protocol, with no WebDriver binary and …
857048active
hect0x7/JMComic-Crawler-Python
A Python library providing an API client for the JMComic (18comic) site, supporting both web and mobile endpoints, with album downloading, …
976997stable
mishushakov/llm-scraper
A TypeScript library that turns any webpage into structured data using LLMs, built on Playwright and the Vercel AI SDK. It supports multipl…
676915active
bda-research/node-crawler
node-crawler is a TypeScript web crawler/spider library for Node.js that fetches pages and provides server-side DOM parsing with automatic …
916799active
VeNoMouS/cloudscraper
A Python library that wraps Requests to bypass Cloudflare's anti-bot protection pages (IUAM), supporting challenge types v1, v2, v3, and Tu…
346726active
adbar/trafilatura
Trafilatura is a Python package and command-line tool for crawling the web and extracting main text, metadata, and comments from raw HTML w…
936709active
davidteather/TikTok-Api
An unofficial Python API wrapper for TikTok.com that retrieves trending content, user information, hashtags, and video data without authent…
926593active
brightdata/cli
The official Bright Data CLI (npm package @brightdata/cli) providing terminal access to Bright Data's web scraping, search, and structured …
816406active
lexiforest/curl_cffi
curl_cffi is a Python binding for a curl-impersonate fork via cffi, providing an HTTP client that can impersonate browser TLS/JA3, HTTP/2, …
996390active
chyroc/WechatSogou
A Python library providing a scraping API for WeChat official accounts based on Sogou WeChat search. It lets you look up account info and r…
546377active
Python3WebSpider/ProxyPool
A self-hosted proxy pool service that periodically scrapes free proxy sites, stores and scores them in Redis, tests their availability, and…
636243active
epiral/bb-browser
bb-browser is a TypeScript CLI and MCP server that lets AI agents control your real Chrome browser using your existing login state, exposin…
706128active
MontFerret/ferret
Ferret is a declarative, expression-oriented query language (FQL) with an embeddable Go runtime for querying, transforming, and automating …
836008active
hect0x7/JMComic-APK
A GitHub Actions-based automation that periodically checks for updates to the JM Comic (18comic) Android APK and publishes new versions as …
935992active
omkarcloud/botasaurus
Botasaurus is an all-in-one Python web scraping framework with built-in anti-detection, caching, parallelization, and proxy support. It let…
725688active
rmax/scrapy-redis
Redis-based components for Scrapy that enable distributed crawling and scraping by sharing a Redis queue across multiple spider instances. …
605642active
gosom/google-maps-scraper
An open-source Go tool that scrapes Google Maps to extract business data such as names, addresses, phone numbers, websites, ratings, review…
925630active
xuejianxianzun/PixivBatchDownloader
A browser extension (Chrome, Edge, Firefox) for batch downloading illustrations, manga, Ugoira animations, and novels from Pixiv. It offers…
985564active
browser-act/skills
BrowserAct Skills is a Python-based browser automation CLI designed for AI agents, providing real-browser control with anti-bot evasion (st…
605446active
AhmadIbrahiim/Website-downloader
A Node.js web application that downloads the complete source code of any website, including all assets like JavaScripts, stylesheets, and i…
765245active
Yuukiy/JavSP
JavSP is a Python command-line tool that scrapes adult video (JAV) metadata from multiple websites, aggregates the data, and generates NFO …
255136active
lc/gau
gau (getallurls) is a Go CLI tool that fetches known URLs for a given domain from AlienVault's Open Threat Exchange, the Wayback Machine, C…
565076active
apify/apify-mcp-server
The Apify MCP Server exposes thousands of Apify Store scrapers, crawlers, and automation tools to AI agents via the Model Context Protocol,…
845059active
jaypyles/Scraperr
Scraperr is a self-hosted web scraping application with a web UI that lets users scrape websites without writing code, using XPath-based ex…
104910active
MechanicalSoup/MechanicalSoup
A Python library for automating interaction with websites, built on Requests and BeautifulSoup. It handles cookies, redirects, link followi…
664888active
bjesus/pipet
Pipet is a command-line web scraper written in Go that extracts data from online assets using HTML parsing, JSON parsing, and client-side J…
244772active
l0o0/translators_CN
A community-maintained collection of Zotero translators for Chinese academic and general websites, enabling Zotero to scrape citation metad…
764721active
DedSecInside/TorBot
TorBot is a Python CLI tool for OSINT on the dark web, crawling .onion sites over the Tor network and building link trees. It can save craw…
974716active
xroche/httrack
HTTrack is a free offline browser utility that recursively downloads websites to a local directory, rewriting links so the mirrored copy ca…
994702stable
ultrafunkamsterdam/nodriver
Nodriver is a fully asynchronous Python browser automation and web scraping library, and the official successor to Undetected-Chromedriver.…
624699active
lecepin/WeChatVideoDownloader
A convenient desktop GUI application for downloading videos from WeChat Channels (WeChat Video Accounts). It intercepts and captures video …
104677active
d60/twikit
Twikit is a free Python library that wraps Twitter's internal API, allowing posting, searching, and scraping tweets without an official API…
604631active
dataabc/weibo-crawler
A Python crawler for Sina Weibo that scrapes user profiles and posts, exporting data to CSV, JSON, MySQL, MongoDB, or SQLite, and optionall…
744625active
Keiyoushi Extensions
A community-maintained repository of extensions (APKs) for Mihon and its forks, providing manga source plugins. The source code for the ext…
714595active
joeyism/linkedin_scraper
A Python library that scrapes LinkedIn for user, company, and job data using Playwright with an async API. It provides Pydantic data models…
824452active
sparklemotion/mechanize
Mechanize is a Ruby library for automating interaction with websites. It handles cookies, redirects, link following, and form submission wh…
914439stable
rachelos/we-mp-rss
A self-hosted WeChat official account (公众号) subscription assistant that scrapes articles, generates RSS feeds, and converts content to Mark…
844396active
UltimaHoarder/UltimaScraper
A Python-based scraper that downloads all media (photos, videos) from OnlyFans accounts using the user's own session authentication. It sto…
234271active
kanasimi/work_crawler
A multi-language downloader application that batch-downloads web novels (converting them to EPUB) and comics from a large list of Chinese, …
664193active
Patchright
Patchright is a patched, undetected fork of the Playwright browser automation framework that evades bot-detection systems like Cloudflare. …
934191active
speedyapply/JobSpy
JobSpy is a Python library that scrapes job postings from popular job boards like LinkedIn, Indeed, Glassdoor, Google, and ZipRecruiter con…
574165active
dotnetcore/DotnetSpider
DotnetSpider is a .NET Standard web crawling and scraping framework that is lightweight, efficient, and cross-platform. It supports distrib…
574138active
Lucksi/Mr.Holmes
Mr.Holmes is a Python-based OSINT (open-source intelligence) CLI tool that gathers information about usernames, domains, phone numbers, and…
534112active
nghuyong/WeiboSpider
A continuously maintained Python web scraping tool for Sina Weibo built on Scrapy and the new weibo.com API. It collects user profiles, pos…
734109active
RipMeApp/ripme
RipMe is a cross-platform Java application that bulk-downloads image albums from websites like Reddit, Imgur, Twitter, Instagram, and Tumbl…
984104active

page 1 / 6 next →