Ross ROSS = Recommend OSS · open-source software intelligence for agents

domain: crawlers

581 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
chaitin/rad
Rad (Radium) is a browser-based web crawler built for security scanning, driving a real Chrome browser to discover URLs and requests across…
231514maintenance
VideoData/DY-Data
A collection of Douyin (Chinese TikTok) scraping tools and API source code covering search, user, video, live stream, comments, danmaku, an…
321510maintenance
fossasia/event-collect
A Python CLI tool that scrapes event website listings (e.g., EventBrite search results) and converts them into the Open Event JSON format. …
101505maintenance
78/ssbc
The source code of the Shousibaocai (手撕包菜) website, a self-hosted BitTorrent magnet link search engine. It combines a Node.js DHT network s…
321476maintenance
AlexCSDev/PatreonDownloader
A command-line application for downloading content posted by creators on patreon.com, including posts, attachments, descriptions, and embed…
691471maintenance
udacimak/udacimak
Udacimak is a command-line tool that downloads Udacity Nanodegree and course content, including videos and materials, and renders them loca…
321457maintenance
liguobao/HouseSearch
A map-driven rental housing aggregation platform that continuously crawls public rental listings from sources like Douban, Beike, and Xiaoh…
761439maintenance
xisuo67/XHS-Spider
XHS-Spider is a polished Windows desktop (WPF, .NET 6) tool for collecting Xiaohongshu (Little Red Book) data, including keyword and user s…
641435maintenance
damklis/DataEngineeringProject
An end-to-end data engineering project that scrapes news from RSS feeds via Airflow-scheduled Python scrapers and streams them through Kafk…
321429maintenance
jonnnnyw/php-phantomjs
A PHP library that wraps the PhantomJS headless browser, letting PHP applications load web pages with full JavaScript support and inspect t…
231426maintenance
kotartemiy/pygooglenews
A Python wrapper around the Google News RSS feed providing top stories, topic and geolocation feeds, and full-text search with date-range a…
321391maintenance
PKUJohnson/OpenData
OpenDataTools is a Python library that scrapes financial and investment data from various websites and exposes it through simple, easy-to-u…
321367maintenance
martinsbalodis/web-scraper-chrome-extension
Web Scraper is a Chrome browser extension for extracting data from web pages without writing code. Users define sitemaps describing how to …
321363maintenance
felipecsl/wombat
Wombat is a lightweight Ruby web crawler and scraper library with an elegant DSL for extracting structured data from web pages. It lets dev…
731360maintenance
ai-to-ai/Auto-Gmail-Creator
A Python Selenium-based bot that bulk-creates Gmail accounts automatically, using sms-activate.org for phone verification and webdriver-man…
311360maintenance
huaying/instagram-crawler
A Python command-line crawler that scrapes Instagram posts, profiles, and hashtag data using Selenium and ChromeDriver, without the officia…
321350maintenance
adriancooney/puppeteer-heap-snapshot
A TypeScript library and CLI for capturing Chrome DevTools heap snapshots from Puppeteer-controlled pages and querying them for objects wit…
231350maintenance
scrapinghub/frontera
Frontera is a Python web crawling framework that implements a scalable crawl frontier, storing and prioritizing links extracted by crawlers…
341332maintenance
laramies/metagoofil
Metagoofil is a Python command-line OSINT tool that searches Google for public documents (pdf, doc, xls, ppt) on target websites, downloads…
321313maintenance
c4tcom/Katana
Katana-ds is a Python CLI tool that automates advanced Google queries known as Google Dorks (Google Dorking), with optional Tor support for…
101304maintenance
InstaPy/instagram-profilecrawl
A Python script from the InstaPy ecosystem that crawls public Instagram profile information such as post counts, follower counts, and post …
321301maintenance
devanshbatham/FavFreak
A Python CLI tool that fetches favicon.ico files from lists of URLs, computes their mmh3 hashes, and groups domains/subdomains/IPs by match…
231298maintenance
sunra/php-simple-html-dom-parser
A PHP library adapting Simple HTML DOM Parser for Composer and PSR-0, allowing easy HTML parsing and manipulation. It supports invalid HTML…
231281maintenance
Sniper970119/dianping_spider
A Python web scraper for Dianping (大众点评) that extracts search results, shop details, and reviews, handling dynamic font encryption without …
721278maintenance
dragnet-org/dragnet
Dragnet is a Python library that uses machine learning models to extract the main article content, and optionally user comments, from HTML …
361274maintenance
useragents/Zefoy-TikTok-Automator
A Python/Selenium automation script that drives zefoy.com to send TikTok followers, views, likes, shares, favorites, and comment likes. It …
321265maintenance
kiddyuchina/Beanbun
Beanbun is a multi-process web crawler framework written in PHP, built on Workerman with Guzzle as the default downloader. It supports dist…
611259maintenance
dchrastil/ScrapedIn
A Python CLI tool that scrapes LinkedIn without API restrictions to enumerate employees of a target company for red team or social engineer…
321235maintenance
istresearch/scrapy-cluster
Scrapy Cluster is a distributed web scraping framework built on Scrapy that uses Redis to coordinate crawl requests and Kafka as a data bus…
101225maintenance
vysecurity/LinkedInt
LinkedInt is a Python CLI tool for LinkedIn reconnaissance that scrapes employee profiles for a target company and generates an HTML report…
101214maintenance
Mapaler/PixivUserBatchDownload
A userscript (browser extension-style script) that batch-downloads all public works of a Pixiv artist directly from the Pixiv website, send…
471191maintenance
timwhitez/crawlergo_x_XRAY
A Python glue script that combines the crawlergo dynamic crawler with the XRAY passive vulnerability scanner, replaying crawled URLs throug…
321182maintenance
yutto-dev/bilili
bilili is a Python CLI tool for downloading Bilibili videos (including bangumi/season content) along with danmaku comments and subtitles. I…
101180maintenance
WebSpiderUtils/verification_code
A research repository documenting approaches and code for solving mainstream CAPTCHA systems such as Geetest, NetEase Yidun, and Aliyun CAP…
321165maintenance
dixudx/tumblr-crawler
A Python script that downloads all photos and videos from specified Tumblr blogs. It supports batch site lists, proxy configuration, and sk…
591157maintenance
holgerd77/django-dynamic-scraper
A Django app that builds on the Scrapy framework and lets you create and manage Scrapy spiders through the Django admin interface. It remov…
231157maintenance
s045pd/DarkNet_ChineseTrading
A Python-based real-time crawler that monitors Chinese-language darknet marketplaces over Tor, with automatic account registration, login, …
101149maintenance
Jinnrry/RobotHelper
RobotHelper is an Android automation script framework written in Java, providing common building blocks like screen capture, image-based po…
231136maintenance
juancarlospaco/faster-than-requests
A Python 3 HTTP client library claiming to be much faster than the popular Requests library, implemented with a compiled core and minimal d…
671129maintenance
bonfy/github-trending
A Python script that scrapes GitHub's trending page daily and archives the most popular repositories as Markdown files. It is designed to r…
771128maintenance
kohlschutter/boilerpipe
boilerpipe is a Java library for removing boilerplate (ads, navigation, headers) from HTML pages and extracting the main full text content.…
321127maintenance
medialab/artoo
artoo.js is a JavaScript library injected into a webpage's context (typically via a bookmarklet) that provides client-side web scraping uti…
231119maintenance
exorde-labs/exorde-client
The Exorde client is a Python CLI worker node for the Exorde Network, a decentralized protocol where participants scrape social media and w…
601085maintenance
acikyazilimagi/afet-org
An open-source earthquake relief platform (depremyardim.com / afetharita.com) that aggregates calls for help from Twitter, WhatsApp, Telegr…
311079maintenance
jonbakerfish/TweetScraper
TweetScraper is a Scrapy-based crawler that scrapes tweets and user information from Twitter Search without using Twitter's official APIs. …
231062maintenance
utkarshkukreti/select.rs
A Rust library for extracting useful data from HTML documents, built around predicate-based DOM queries. It is designed for web scraping ta…
381020maintenance
Algebra-FUN/WeReadScan
A Python library that uses Selenium headless browsers to scan purchased books from WeRead (WeChat Reading) and convert them into local PDF …
321002maintenance
JSREI/ast-hook-for-js-RE
A browser memory roaming tool for JavaScript reverse engineering that hooks variable assignments via AST-transformed proxy responses. It le…
231912experimental
tinyfish-io/bigset-oss
BigSet is a self-hostable application that turns a natural-language sentence into a structured, regularly refreshed dataset by dispatching …
761684experimental
kkyon/botflow
Botflow is a Python dataflow programming framework for building data pipelines using pipes and routes, with parallelism via coroutines and …
621196experimental
oxylabs/ai-scraper-py
AI-Scraper is a Python library and scrape agent from Oxylabs AI Studio that extracts data from webpages using natural language prompts inst…
511055experimental
twintproject/twint
Twint is a Python CLI tool and library that scrapes tweets, followers, following, and likes from Twitter without using the official API or …
1016398abandoned
wechat-article/wechat-article-exporter
An online batch downloader for WeChat Official Account articles that exports posts in HTML, JSON, Excel, TXT, Markdown, and DOCX formats, i…
6812789abandoned
xiandanin/magnetW
MagnetW is a cross-platform desktop application built with Electron and Vue that aggregates magnet link search results from multiple torren…
1011271abandoned
yangyangwithgnu/hardseed
hardseed is a C++ command-line tool that scrapes adult forum threads (aicheng, caoliu) to batch-download images and torrent seed files, wit…
329198abandoned
TonyChen56/WeChatRobot
A C++ WeChat hooking toolkit and robot framework providing wxhook APIs, WeChat database decryption, and Official Account (公众号) scraping, pl…
587188abandoned
xchaoinfo/fuck-login
A Python library that simulates programmatic login to popular Chinese websites like Zhihu, Weibo, Baidu, and Douban, built on requests, Pil…
105870abandoned
hanc00l/wooyun_public
A crawler and search application for the archived Wooyun.org security vulnerability disclosure platform, containing ~40k-88k public vulnera…
104399abandoned
qiyeboy/IPProxyPool
IPProxyPool is a Python proxy pool service that crawls free proxy IPs from the web, validates them, stores them in a database (SQLite by de…
234286abandoned
bisguzar/twitter-scraper
A Python library that scrapes Twitter's frontend JavaScript API without authentication, letting users fetch tweets from profiles or hashtag…
104005abandoned
pyppeteer/pyppeteer
Pyppeteer is an unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. It supports page navigati…
323942abandoned
miyakogi/pyppeteer
An unofficial Python port of Puppeteer for automating headless Chrome/Chromium browsers via asyncio. This original repository has moved to …
103550abandoned
bowenpay/wechat-spider
A Python-based web crawler for scraping articles from WeChat public accounts (微信公众号), built on Django with MySQL and Redis, including a web…
323369abandoned
LiuXingMing/SinaSpider
A Python web crawler for Sina Weibo (Chinese microblog) built on Scrapy, with three versions: a standalone spider, a distributed version us…
323285abandoned
gnemoug/distribute_crawler
A distributed web crawler built on Scrapy, Redis, MongoDB, and Graphite, demonstrated with a spider for a Chinese book-download site. Redis…
323238abandoned
harismuneer/Ultimate-Social-Scrapers
A collection of Python-based scraping tools that extract public data from Facebook, Instagram, and Twitter (X), including posts, media, fol…
443151abandoned
airingursb/bilibili-user
A Python web crawler that scrapes Bilibili user profiles (id, nickname, gender, avatar, level, birthday, location, etc.) and stores them in…
323090abandoned
NikolaiT/GoogleScraper
GoogleScraper is a Python module and CLI tool for scraping search engine results from Google, Bing, Yandex, DuckDuckGo and others, with sup…
322874abandoned
YahooArchive/anthelion
Anthelion is an Apache Nutch plugin for focused crawling of semantic data embedded in HTML pages. It uses an online learning classifier to …
102827abandoned
lanbing510/DouBanSpider
A Python web scraper for Douban Books that crawls book listings by tag, storing ratings and review counts into Excel files. The author also…
322786abandoned
loadchange/amemv-crawler
A Python 3 script that downloads all videos from a specified Douyin (TikTok China) user account, as well as all videos under a given challe…
322641abandoned
taspinar/twitterscraper
A Python library that scrapes tweets and user information from Twitter using requests and BeautifulSoup, without relying on Twitter's offic…
232461abandoned
scrapoxy/scrapoxy
Scrapoxy was an open-source proxy manager for web scraping that aggregated proxies from cloud providers and other sources behind a single A…
622414abandoned
egrcc/zhihu-python
A Python 2.7 library for scraping content from Zhihu, a Chinese Q&A platform, including questions, answers, users, and favorites. It can ex…
322335abandoned
chiphuyen/lazynlp
A Python library for crawling, cleaning, and deduplicating web pages to build massive monolingual text datasets, suitable for training lang…
232284abandoned
PaulMcInnis/JobFunnel
JobFunnel is a Python CLI tool that scrapes job postings from multiple job websites (Indeed, Glassdoor, LinkedIn) into a single deduplicate…
102180abandoned
minimaxir/facebook-page-post-scraper
A Python script collection that scrapes all posts, reactions, and comments from public Facebook Pages and open Groups via the Facebook Grap…
102135abandoned
simplecrawler/simplecrawler
simplecrawler is a flexible, event-driven web crawler library for Node.js with a configurable queue system, robots.txt support, and link di…
102134abandoned
althonos/InstaLooter
InstaLooter is a Python CLI tool that downloads pictures and videos from Instagram profiles without using the official API. It is a re-impl…
232098abandoned
cycz/jdBuyMask
A Python automation tool that monitored JD.com (Jingdong) for face mask stock during the COVID-19 pandemic and automatically placed purchas…
321849abandoned
hu17889/go_spider
go_spider is a concurrent web crawler framework written in Go, designed for crawling vertical communities with a flexible, modular architec…
231818abandoned
node-js-libs/node.io
node.io is a Node.js web scraping and data extraction library originally written in 2010. It is explicitly no longer maintained, with the a…
321792abandoned
bughandler/cnki-downloader
A small desktop tool for searching and downloading academic literature from CNKI (China National Knowledge Infrastructure). Its backend int…
481762abandoned
ZFC-Digital/puppeteer-real-browser
A Node.js library that wraps Puppeteer with a real-browser profile to bypass bot detection systems like Cloudflare and Turnstile captchas. …
351642abandoned
erma0/douyin
A Python crawler for Douyin (Chinese TikTok) that collected public data such as account profiles, likes, favorites, music, hashtags, search…
761602abandoned
fossasia/loklak_wok_android
Loklak Wok is an Android app that acts as a harvesting peer for the loklak_server, collecting social media messages (tweets) and pushing th…
101558abandoned
GravityLabs/goose
Goose is a Scala library (originally Java) that extracts the main body text, metadata, publish date, embedded videos, and top image from ne…
101526abandoned
yhat/scrape
A Go library providing a higher-level interface over golang.org/x/net/html for web scraping. It offers generic tree traversal helpers like …
101515abandoned
qinxuye/cola
Cola is a high-level distributed crawling framework in Python for scraping pages and extracting structured data from websites. The same cra…
101499abandoned
johntitus/node-horseman
A Node.js library providing a chainable, Promise-based API for controlling the PhantomJS headless browser, supporting page navigation, form…
321485abandoned
fossasia/loklak_scraper_js
A collection of JavaScript scrapers for the loklak project that extract data from websites like Twitter and Quora and output JSON resemblin…
101473abandoned
keenwon/antcolony
AntColony is a Node.js-based BitTorrent DHT network crawler that collects active infohashes, downloads and parses torrent files, and stores…
321456abandoned
jamesturk/scrapeghost
scrapeghost is an experimental Python library that uses OpenAI's GPT models to scrape structured data from websites without writing page-sp…
491442abandoned
OpnTec/parliament-scraper
A collection of scrapers (in Python, Ruby, and Scala) that download public parliamentary data such as written questions from the EU Parliam…
321403abandoned
leonardocardoso/SwiftLinkPreview
A Swift library that generates link previews from URLs by extracting titles, relevant text, and images. It supports iOS, macOS, watchOS, an…
101385abandoned
dinubs/jam-api
Jam API is a hosted service (self-hostable Node.js app) that turns any website into a JSON API by extracting data using CSS query selectors…
321363abandoned
Vespa314/bilibili-api
A collection of Bilibili (B站) API documentation and Python tooling for scraping video, user, comment, danmaku, and bangumi data. It include…
321357abandoned
Jefferson-Henrique/GetOldTweets-python
A Python library that retrieves old tweets by mimicking the JSON calls Twitter Search makes in the browser, bypassing the official API's ti…
321340abandoned
adyzng/jd-autobuy
A Python 2.7 command-line scraper that logs into JD.com (via QR code or credentials), monitors product stock and price, and automatically p…
321309abandoned
LiuRoy/zhihu_spider
A Scrapy-based web crawler that scrapes Zhihu user profiles and their follower/following relationship graphs, storing data in MongoDB. It s…
321282abandoned

← prev page 5 / 6 next →