Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: crawlers

581 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
scrapinghub/splash
Splash is a lightweight, scriptable headless browser exposed as a service with an HTTP API, implemented in Python 3 using Twisted and Qt5. …
324187maintenance
DataDog/guarddog
GuardDog is a CLI tool from Datadog that identifies malicious packages on PyPI, npm, Go modules, Rust crates, RubyGems, GitHub Actions, and…
991194active
constverum/ProxyBroker
ProxyBroker is an asynchronous Python tool that finds public HTTP(S) and SOCKS4/5 proxies from ~50 sources and concurrently checks their ty…
324159maintenance
xenova/chat-downloader
Chat Downloader is a Python tool and library for retrieving chat messages from livestreams, videos, clips, and past broadcasts on platforms…
451189active
intoli/user-agents
A JavaScript/TypeScript npm package for generating random user agents weighted by real-world market share, with daily-updated data. It also…
771188active
goodreasonai/ScrapeServ
ScrapeServ is a self-hosted API service that accepts a URL and returns the website's data along with browser screenshots, using Playwright …
251181active
rchipka/node-osmosis
Osmosis is an HTML/XML parser and web scraper library for Node.js built on native libxml C bindings. It offers a chainable, promise-like in…
324107maintenance
grangier/python-goose
Python-Goose is a Python library that extracts the main body text, metadata, top image, and embedded videos from news article web pages. It…
644106maintenance
online-judge-tools/oj
A command-line tool that automates solving problems on online judges like AtCoder, Codeforces, and HackerRank. It downloads sample and syst…
231169active
zu1k/proxypool
A Go service that automatically crawls proxy nodes (ss, ssr, vmess, trojan) from Telegram channels, subscription URLs, and the public inter…
234027maintenance
fanpei91/torsniff
torsniff is a Go CLI tool that sniffs torrent metadata from the BitTorrent network by participating in the DHT and connecting to peers to d…
104014maintenance
tholian-network/stealth
Stealth is a secure, privacy-focused web browser, scraper, and proxy built in JavaScript that emphasizes automation, bandwidth efficiency, …
321145active
datawhores/OF-Scraper
OF-Scraper is a command-line tool for downloading media from OnlyFans and performing bulk actions like liking or unliking posts. It is a re…
801143active
AndyTheFactory/newspaper4k
Newspaper4k is a Python library and CLI for scraping and curating news articles, extracting text, titles, authors, publish dates, and metad…
841140active
fwonggh/Bthub
Bthub is a magnet link and torrent search engine, and this repository serves as its official address release page listing current and backu…
761133active
webrecorder/browsertrix-crawler
Browsertrix Crawler is a standalone browser-based high-fidelity web crawling system that runs in a single Docker container. It uses Puppete…
991120active
platonai/Browser4
Browser4 is an AI-native browser engine built in Kotlin for autonomous agents, intelligent data extraction, and large-scale web automation.…
1001114active
elixir-crawly/crawly
Crawly is a high-level web crawling and scraping framework for Elixir, modeled after Scrapy, where developers define spiders that fetch pag…
371114active
nottelabs/reverse-api-engineer
Reverse API Engineer is a Python CLI tool that captures browser network traffic (HAR) from a website and uses a configured AI model to gene…
841113active
bellingcat/auto-archiver
A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supp…
921109active
vifreefly/kimuraframework
Kimuraframework (Kimurai) is a Ruby web scraping framework with an AI-assisted DSL: an LLM generates XPath selectors from a schema on first…
611102active
GPTaku Plugins
insane-search is a Claude Code plugin that reads public web pages that would otherwise be blocked (403, CAPTCHA, WAF), escalating through p…
601102active
Tyrrrz/YoutubeExplode
A .NET library providing an abstraction layer over YouTube's internal API to query metadata for videos, playlists, and channels, and to res…
993717maintenance
ruipgil/scraperjs
Scraperjs is a Node.js web scraping library offering two scrapers: a lightweight StaticScraper using cheerio for static HTML, and a Dynamic…
323714maintenance
techtanic/Discounted-Udemy-Course-Enroller
A Python application (with GUI and CLI variants) that scrapes websites for 100% off Udemy course coupons and automatically enrolls the user…
711081active
jmcarp/robobrowser
RoboBrowser is a Pythonic library for browsing the web without a standalone browser, combining Requests for HTTP sessions with BeautifulSou…
323692maintenance
jae-jae/fetcher-mcp
A Model Context Protocol (MCP) server that fetches web page content using a Playwright headless browser, executing JavaScript to handle dyn…
481076active
1061700625/WeChat_Article
A PyQt5 desktop application that crawls and downloads all articles from a specified WeChat official account. It uses Selenium to log in and…
621075active
scrapfly/scrapfly-scrapers
A collection of educational Python web scraping scripts for over 40 popular domains such as Amazon, AliExpress, BestBuy, and Twitter, built…
741074active
soxoj/socid-extractor
socid_extractor is a Python library and CLI that extracts structured account metadata and stable internal identifiers (usernames, UIDs, GAI…
921073active
tomnomnom/assetfinder
A Go command-line tool that discovers domains and subdomains potentially related to a given domain by querying multiple passive sources lik…
233666maintenance
unitedstates/congress
A community-run Python toolkit that collects and converts official U.S. Congress data—bills, amendments, roll call votes, nominations, and …
521060active
cantino/selectorgadget
SelectorGadget is an open-source bookmarklet and Chrome extension that generates CSS selectors for page elements through point-and-click se…
691059active
spider-ios/autox-release
AutoX is a desktop social media operations tool that automates one-click publishing of videos to multiple platforms such as Douyin, TikTok,…
651059active
Henryhaohao/Bilibili_video_download
A Python tool for downloading videos from Bilibili, supporting single-part and multi-part (分P) videos, bangumi episodes, and multiple downl…
323583maintenance
jaimeiniesta/metainspector
MetaInspector is a Ruby gem for web scraping that fetches a given URL and exposes its title, meta description, keywords, links, images, cha…
691049active
caolvchong-top/twitter_download
A Python command-line tool that scrapes and downloads images, videos (including GIFs), and text from Twitter/X user timelines. It supports …
751045active
mandatoryprogrammer/thermoptic
Thermoptic is a stealth HTTP proxy that routes requests from any HTTP client (like curl) through a containerized Chrome instance so the tra…
511045active
turicas/brasil.io
The backend of Brasil.IO, a platform that collects, cleans, and publishes Brazilian public open datasets in accessible formats. It automate…
771044active
tmwgsicp/wechat-download-api
An open-source API service for fetching WeChat official account articles, generating standard RSS 2.0 feeds, and exporting entire account a…
771043active
lm-rebooter/NuggetsBooklet
A Node.js script that downloads Juejin (掘金) booklet content for personal study by reusing the user's authenticated browser cookies. It expo…
661041active
Anorov/cloudflare-scrape
A Python module (cfscrape) built on Requests that bypasses Cloudflare's JavaScript anti-bot challenge page ('I'm Under Attack Mode') so scr…
233538maintenance
jfilter/clean-text
A Python package for cleaning and normalizing messy text, especially user-generated content from the web and social media. It fixes unicode…
811027active
carcabot/tiktok-signature
A self-hosted Node.js service that generates valid X-Bogus and X-Gnarly signature tokens for TikTok API requests using a headless browser r…
931025active
wnma3mz/wechat_articles_spider
A Python library for scraping WeChat Official Account articles, including article URLs, reading counts, likes, and comments. It can also do…
233481maintenance
daijro/hrequests
hrequests is a Python HTTP client library that replaces the requests library with browser TLS fingerprint replication, HTTP/2 support, and …
301023active
JosephLai241/URS
URS (Universal Reddit Scraper) is a comprehensive command-line tool written in Python (with Rust components) for scraping and archiving Red…
641020active
owner888/phpspider
phpspider is a PHP web crawling framework that lets developers build scrapers with a simple config array, handling multi-process workers, l…
233462maintenance
wujunwei928/parse-video
A Go library and CLI tool that parses short-video share links from 25+ Chinese platforms (Douyin, Kuaishou, Bilibili, Xiaohongshu, Weibo, e…
811018active
Vinyzu/Botright
Botright is a Python browser automation framework built on Playwright that provides undetectable, fingerprint-changing stealth browsing. It…
761017active
pea3nut/Pxer
Pxer is a userscript (installed via Tampermonkey) that acts as a crawler for pixiv.net, letting users batch-fetch artworks, collections, an…
271010active
wreq
wreq is an ergonomic, privacy-aware HTTP client written in Rust with a Python binding (wreq-python) that provides high-fidelity browser TLS…
901006active
ranahaani/GNews
GNews is a lightweight Python package that queries the Google News RSS feed and returns article results as usable JSON. It supports keyword…
681006active
JoMingyu/google-play-scraper
A Python library that provides APIs to crawl the Google Play Store for app details, reviews, and other data without any external dependenci…
321006active
ma6254/FictionDown
FictionDown is a Go-based command-line tool for batch downloading and crawling web novels from sites like Qidian and Biquge. It supports mu…
241006active
jackwener/wechat-article-to-markdown
A Python CLI tool that fetches WeChat Official Account articles using anti-detection browser automation (Camoufox) and converts them to cle…
481002active
kevinzg/facebook-scraper
A Python library for scraping public Facebook pages, groups, profiles, and posts without requiring an API key. It provides a simple get_pos…
323270maintenance
wenbochang888/house
A Java web scraper built with SpringBoot, HttpClient, and JSoup that crawls famous Tianya forum threads about China's housing market and co…
323231maintenance
scrapy-plugins/scrapy-splash
A Scrapy plugin that integrates the Splash headless browser service to enable crawling and scraping of JavaScript-rendered web pages. It pr…
263227maintenance
ferventdesert/Hawk
Hawk is a visual crawler and ETL IDE written in C#/WPF that lets users graphically scrape webpages, clean, transform, and store data withou…
233213maintenance
CrawlScript/WebCollector
WebCollector is an open-source Java web crawler framework that provides simple interfaces for building multi-threaded web crawlers quickly.…
623083maintenance
jaeles-project/gospider
GoSpider is a fast web spider/crawler written in Go that crawls sites in parallel and extracts URLs from sitemaps, robots.txt, JavaScript f…
232993maintenance
kotartemiy/newscatcher
A Python package that programmatically collects normalized news articles from thousands of news websites, filterable by topic, country, and…
322987maintenance
Threezh1/JSFinder
JSFinder is a Python command-line tool that crawls a website's JavaScript files and extracts URLs and subdomains using regex parsing. It su…
322976maintenance
facundoolano/google-play-scraper
A Node.js library that scrapes application data from the Google Play store, exposing methods for app details, search, reviews, permissions,…
662949maintenance
howie6879/owllook
owllook is a self-hosted vertical search engine for Chinese web novels, built on Python with Sanic, MongoDB, and Redis. It aggregates resul…
322871maintenance
CharlesPikachu/DecryptLogin
A Python library providing programmatic login APIs for popular websites (Weibo, Bilibili, Zhihu, GitHub, Taobao, etc.) built on the request…
322855maintenance
yann-shi/dht
A Go library implementing the BitTorrent DHT protocol (BEP-3, 5, 9, 10) with two modes: a standard DHT server and a crawling mode for harve…
232769maintenance
brianway/webporter
webporter is a Java crawler application built on the webmagic framework that demonstrates a complete pipeline of data crawling, persistence…
232766maintenance
DormyMo/SpiderKeeper
SpiderKeeper is a self-hosted web-based admin dashboard for managing Scrapy spiders running on Scrapyd servers. It provides spider scheduli…
322763maintenance
jeanphix/Ghost.py
Ghost.py is a scriptable WebKit-based web client library for Python, built on PySide2/Qt5, allowing programmatic page loading and content i…
322756maintenance
luin/readability
A Node.js library that extracts clean, readable article content from any web page, based on arc90's readability project. It returns the art…
322519maintenance
xtuhcy/gecco
Gecco is a lightweight, easy-to-use web crawler framework for Java that lets developers define crawlers with annotation-based jQuery-style …
512510maintenance
lorien/grab
Grab is a Python web scraping framework providing HTTP request handling, proxy/cookie support, and XPath-based HTML parsing, plus a Spider …
512463maintenance
decaywood/XueQiuSuperSpider
A Java 8 web scraping framework for collecting stock data from Xueqiu (Snowball) and other Chinese financial sites. It is built around comp…
322433maintenance
paquettg/php-html-parser
A PHP library that parses HTML into a DOM and lets you find and manipulate tags using CSS selectors, similar to jQuery. It is designed for …
322398maintenance
QianyanTech/Image-Downloader
A Python application that crawls and downloads images from Google, Bing, and Baidu using Selenium or API drivers. It offers both a PyQt5 GU…
232362maintenance
lucasjinreal/weibo_terminater
A Python-based web scraper that crawls Weibo (Sina's microblog platform) to collect user posts, comments, followers, and conversation pairs…
322317maintenance
ageitgey/node-unfluff
A Node.js library and CLI tool that automatically extracts the main body content and metadata (title, author, date, images, tags, links) fr…
322158maintenance
t9tio/cloudquery
CloudQuery is a tool that turns any website into a JSON API by fetching pages with headless Chrome and extracting data via CSS selectors. I…
322149maintenance
php-embed/Embed
A PHP library that extracts metadata and embed information from any web page or web service using oEmbed, OpenGraph, Twitter Cards, and HTM…
932140maintenance
anouarbensaad/vulnx
VulnX is a Python CLI tool that detects CMS types (WordPress, Joomla, Drupal, etc.), gathers target information like subdomains and DNS rec…
232138maintenance
PuerkitoBio/gocrawl
gocrawl is a polite, slim and concurrent web crawler library written in Go. It respects robots.txt rules, applies per-host crawl delays, an…
232052maintenance
Nekmo/dirhunt
Dirhunt is a Python CLI web crawler optimized for finding and analyzing web directories without brute-forcing paths. It detects 'index of' …
232007maintenance
awolfly9/IPProxyTool
A Python/Scrapy application that crawls free proxy websites, validates the collected proxy IPs against target sites, and stores usable prox…
321998maintenance
Xyntax/POC-T
POC-T is a Python 2.7 plugin-based concurrent framework for penetration testing tasks such as crawling, bruteforcing, and batch PoC/EXP ver…
231936maintenance
scrapy/scrapely
Scrapely is a pure-Python library for extracting structured data from HTML pages. It learns a parser from example pages annotated with the …
321883maintenance
xianhu/PSpider
PSpider is a simple, easy-to-read web spider framework written in Python 3.8+. It uses a Fetcher/Parser/Saver pipeline with queues for mult…
321835maintenance
reworkd/tarsier
Tarsier is a Python library providing vision utilities for LLM-driven web interaction agents. It visually tags interactable page elements w…
171761maintenance
howie6879/ruia
Ruia is an async web scraping micro-framework for Python 3.6+ built on asyncio and aiohttp. It offers declarative Item/Field extraction (XP…
231738maintenance
henson/proxypool
A Golang IP proxy pool service that scrapes free proxy sources, validates them, stores them in a database, and exposes a JSON API for crawl…
321701maintenance
YoongiKim/AutoCrawler
A Python multiprocess image web crawler that downloads images from Google and Naver image search using Selenium and ChromeDriver. It is des…
321691maintenance
th3unkn0n/TeleGram-Scraper
A Python command-line tool that scrapes Telegram groups and exports all member information to CSV using the Telegram API. It also includes …
101672maintenance
aivarsk/scrapy-proxies
A Scrapy downloader middleware that routes requests through random proxies from a configurable list to avoid IP bans. It supports multiple …
321667maintenance
gigablast/open-source-search-engine
Gigablast is a distributed open source web and enterprise search engine with a built-in spider/crawler, written in C/C++ for Linux. It powe…
321601maintenance
propublica/upton
Upton is a Ruby framework that handles the repetitive parts of writing web scrapers, letting developers focus on site-specific CSS selector…
321597maintenance
0xHJK/dumpall
dumpall is a Python command-line tool for exploiting information disclosure vulnerabilities on web servers. It reconstructs source code fro…
231579maintenance
lqqyt2423/wechat_spider
A Node.js WeChat crawler that uses a man-in-the-middle proxy (AnyProxy) to batch-collect WeChat official account article data, including co…
761574maintenance
github/lightcrawler
Lightcrawler is a Node.js CLI tool that crawls a website by following links and runs each discovered page through Google Lighthouse audits.…
101564maintenance
headzoo/surf
Surf is a Go library that implements a stateful virtual web browser controlled programmatically. It supports cookies, history, bookmarks, u…
231544maintenance

← prev page 4 / 6 next →