Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
microsoft/Webwright
Webwright is a Python framework from Microsoft that turns coding LLMs into state-of-the-art browser agents by giving them a terminal to lau…
575947active
lanmaster53/recon-ng
Recon-ng is a full-featured, modular reconnaissance framework for conducting web-based open source intelligence (OSINT) gathering. It offer…
325871active
HapeLee/legado-with-MD3
Legado with MD3 is a free, open-source Android e-book and content aggregation reader, rebuilt from the Legado (阅读 3.0) project with a Mater…
825860active
RedSiege/EyeWitness
EyeWitness is a Python CLI tool that takes screenshots of websites using headless Chromium, captures server header information, and identif…
505829active
joeseesun/qiaomu-anything-to-notebooklm
A Claude Code Skill that ingests content from 15+ sources (WeChat articles, web pages, YouTube, PDFs, EPUB, Office docs, audio) and uploads…
625824active
vantage-sh/ec2instances.info
EC2Instances.info is an open-source web application for comparing Amazon EC2, RDS, and ElastiCache instance types, specs, and pricing acros…
775759active
CharlesPikachu/musicdl
A lightweight music downloader written in pure Python that supports dozens of music and audiobook platforms including NetEase Cloud Music, …
755757active
omkarcloud/botasaurus
Botasaurus is an all-in-one Python web scraping framework with built-in anti-detection, caching, parallelization, and proxy support. It let…
725688active
vikiboss/60s
A collection of free, open-source REST APIs delivering daily news digests, Chinese social media trending lists (Weibo, Bilibili, Douyin, Zh…
925678active
qiye45/wechatVideoDownload
A desktop GUI tool for downloading WeChat Channels (视频号) videos, live streams, live replays, and images. It automatically monitors WeChat t…
715667active
TechXueXi
TechXueXi is an open-source Python automation tool that automatically completes daily study tasks and quizzes on China's Xuexi Qiangguo (学习…
565651active
rmax/scrapy-redis
Redis-based components for Scrapy that enable distributed crawling and scraping by sharing a Redis queue across multiple spider instances. …
605642active
executeautomation/mcp-playwright
A Model Context Protocol (MCP) server that exposes Playwright browser automation to LLM clients like Claude Desktop, Cline, and Cursor IDE.…
475635active
gosom/google-maps-scraper
An open-source Go tool that scrapes Google Maps to extract business data such as names, addresses, phone numbers, websites, ratings, review…
925630active
sallar/github-contributions-chart
A web application that generates a shareable image of a user's entire GitHub contribution history, with multiple color themes. It works aro…
355604active
potree/potree
Potree is a free open-source WebGL-based renderer for viewing large point clouds directly in web browsers. It uses octree-based level-of-de…
505585active
bilawalsidhu/gods-eye-view
God's Eye View is a browser-based photorealistic 3D globe application that visualizes live public data feeds — aircraft, ships, satellites,…
585574active
xuejianxianzun/PixivBatchDownloader
A browser extension (Chrome, Edge, Firefox) for batch downloading illustrations, manga, Ugoira animations, and novels from Pixiv. It offers…
985564active
AngleSharp/AngleSharp
AngleSharp is a .NET library that parses HTML5, SVG, MathML, XML, and CSS into a fully standards-compliant W3C DOM. It supports querySelect…
995531active
tangyoha/telegram_media_downloader
A cross-platform Python application that downloads media files (video, audio, photos, documents) from Telegram chats, channels, and private…
535495active
tebelorg/RPA-Python
A Python package for robotic process automation (RPA) that wraps TagUI to automate web pages, desktop apps, and visual elements via a simpl…
655493active
browser-act/skills
BrowserAct Skills is a Python-based browser automation CLI designed for AI agents, providing real-browser control with anti-bot evasion (st…
605446active
jef/streetmerchant
streetmerchant is a Node.js application that continuously checks retail websites for product stock availability, with optional add-to-cart …
645402active
Neet-Nestor/Telegram-Media-Downloader
A browser userscript (installable via Tampermonkey, Violentmonkey, etc.) that adds media download buttons to the Telegram webapp. It enable…
645253active
AhmadIbrahiim/Website-downloader
A Node.js web application that downloads the complete source code of any website, including all assets like JavaScripts, stylesheets, and i…
765245active
spatie/browsershot
A PHP library that converts web pages or raw HTML into images, PDFs, or rendered HTML strings using Puppeteer and headless Chrome. It suppo…
915241active
Yuukiy/JavSP
JavSP is a Python command-line tool that scrapes adult video (JAV) metadata from multiple websites, aggregates the data, and generates NFO …
255136active
scinfu/SwiftSoup
SwiftSoup is a pure Swift HTML parser library that conforms to the WHATWG HTML5 specification and offers DOM traversal, CSS selectors, and …
995119stable
hakluke/hakrawler
Hakrawler is a fast command-line web crawler written in Go, built on the Gocolly library, that discovers URLs and JavaScript file locations…
665116active
obsidianmd/obsidian-clipper
Obsidian Web Clipper is the official browser extension for Obsidian that lets users highlight web pages and capture content as durable Mark…
835092active
lc/gau
gau (getallurls) is a Go CLI tool that fetches known URLs for a given domain from AlienVault's Open Threat Exchange, the Wayback Machine, C…
565076active
apify/apify-mcp-server
The Apify MCP Server exposes thousands of Apify Store scrapers, crawlers, and automation tools to AI agents via the Model Context Protocol,…
845059active
binbyu/Reader
A lightweight, free, open-source Windows (win32) reader application for txt and epub files, with support for reading online novels via conf…
655057active
FellouAI/eko
Eko is a production-ready JavaScript framework for building reliable AI agents and agentic workflows from natural language, running in both…
634951active
FxEmbed/FxEmbed
FxEmbed is a Cloudflare Worker service that powers FxTwitter, FixupX, and FxBluesky, rewriting X/Twitter and Bluesky links into rich embeds…
774948active
exa-labs/exa-mcp-server
An open-source MCP (Model Context Protocol) server that connects AI assistants and agents to Exa's web search, code search, content fetchin…
664926active
jaypyles/Scraperr
Scraperr is a self-hosted web scraping application with a web UI that lets users scrape websites without writing code, using XPath-based ex…
104910active
MechanicalSoup/MechanicalSoup
A Python library for automating interaction with websites, built on Requests and BeautifulSoup. It handles cookies, redirects, link followi…
664888active
UndeadSec/SocialFish
SocialFish is a Python-based phishing toolkit that clones modern login pages using Playwright browser automation and captures credentials, …
764851active
fb55/htmlparser2
htmlparser2 is a fast, forgiving HTML and XML parser for JavaScript and TypeScript, offering a low-allocation callback interface as well as…
874788stable
davidarroyo1234/InstagramUnfollowers
A browser-based script that scans your Instagram account to identify users who don't follow you back, letting you selectively unfollow them…
754781active
bjesus/pipet
Pipet is a command-line web scraper written in Go that extracts data from online assets using HTML parsing, JSON parsing, and client-side J…
244772active
Integuru-AI/Integuru
Integuru is an AI agent that reverse-engineers platforms' internal APIs by analyzing browser network requests (HAR files) and building depe…
624760active
l0o0/translators_CN
A community-maintained collection of Zotero translators for Chinese academic and general websites, enabling Zotero to scrape citation metad…
764721active
88lin/video_vip
A Tampermonkey/Greasemonkey userscript that integrates multiple third-party parsing interfaces to bypass VIP membership restrictions on Chi…
634718active
DedSecInside/TorBot
TorBot is a Python CLI tool for OSINT on the dark web, crawling .onion sites over the Tor network and building link trees. It can save craw…
974716active
201206030/novel-plus
novel-plus is a full-featured novel/fiction CMS built on Spring Boot, comprising a reader-facing portal, author backend, admin dashboard, a…
884709active
xroche/httrack
HTTrack is a free offline browser utility that recursively downloads websites to a local directory, rewriting links so the mirrored copy ca…
994702stable
ultrafunkamsterdam/nodriver
Nodriver is a fully asynchronous Python browser automation and web scraping library, and the official successor to Undetected-Chromedriver.…
624699active
lecepin/WeChatVideoDownloader
A convenient desktop GUI application for downloading videos from WeChat Channels (WeChat Video Accounts). It intercepts and captures video …
104677active
danburzo/percollate
Percollate is a Node.js command-line tool that converts web pages into readable, well-formatted PDF, EPUB, HTML, or Markdown documents. It …
534667active
ShadowHackrs/gmail-account-creator
A Python-based automation tool for bulk creation of Gmail accounts, featuring anti-detection techniques like human-like typing simulation a…
564657active
d60/twikit
Twikit is a free Python library that wraps Twitter's internal API, allowing posting, searching, and scraping tweets without an official API…
604631active
dataabc/weibo-crawler
A Python crawler for Sina Weibo that scrapes user profiles and posts, exporting data to CSV, JSON, MySQL, MongoDB, or SQLite, and optionall…
744625active
Keiyoushi Extensions
A community-maintained repository of extensions (APKs) for Mihon and its forks, providing manga source plugins. The source code for the ext…
714595active
linsomniac/spotify_to_ytmusic
A set of Python scripts and a GUI for copying liked songs and playlists from Spotify to YouTube Music. It uses the Spotify backup data and …
634582active
syhyz1990/baiduyun
A free open-source Tampermonkey userscript that extracts real direct download links from Chinese cloud storage services (Baidu Netdisk, Ali…
324505active
limbopro/Adblock4limbo
A userscript-based ad-blocking project that removes popups, banners, and video ads on specific streaming, comic, novel, and adult sites via…
774491active
sensepost/gowitness
gowitness is a Go-based command-line utility that uses Chrome Headless to take screenshots of websites, supporting scans of URL lists, CIDR…
754488active
6dylan6/jdpro
A collection of JavaScript automation scripts for the Qinglong panel, primarily for running scheduled JD (Jingdong) sign-in and reward task…
734482active
metatube-community/jellyfin-plugin-metatube
A metadata provider plugin for Jellyfin and Emby media servers that fetches movie and actor metadata from various internet providers via th…
614454active
joeyism/linkedin_scraper
A Python library that scrapes LinkedIn for user, company, and job data using Playwright with an async API. It provides Pydantic data models…
824452active
sparklemotion/mechanize
Mechanize is a Ruby library for automating interaction with websites. It handles cookies, redirects, link following, and form submission wh…
914439stable
GerbenJavado/LinkFinder
LinkFinder is a Python CLI script that discovers endpoints and their parameters in JavaScript files using jsbeautifier and regular expressi…
324439stable
rachelos/we-mp-rss
A self-hosted WeChat official account (公众号) subscription assistant that scrapes articles, generates RSS feeds, and converts content to Mark…
844396active
VonChange/utao
Utao TV (油桃TV, now upgraded to 土拨鼠浏览器) is a third-party browser designed for Android TV boxes and smart TVs that lets users watch live CCTV…
674340active
QasimWani/LeetHub
LeetHub is a browser extension that automatically pushes your LeetCode solutions to GitHub whenever you pass all tests on a problem. It sup…
724339active
anasty17/mirror-leech-telegram-bot
A Python Telegram bot that mirrors or leeches files from direct links, torrents, NZB/Usenet, Google Drive, rclone clouds, and yt-dlp/JDownl…
774279active
UltimaHoarder/UltimaScraper
A Python-based scraper that downloads all media (photos, videos) from OnlyFans accounts using the user's own session authentication. It sto…
234271active
wasi-master/13ft
A self-hosted web service that bypasses paywalls on news sites by fetching pages as GoogleBot, a replacement for the defunct 12ft.io. It se…
834266active
hoothin/UserScripts
A collection of Greasemonkey/Tampermonkey userscripts by hoothin, including Pagetual (auto-pager infinite scrolling), Picviewer CE+ (online…
764264active
codeceptjs/CodeceptJS
CodeceptJS is a Node.js end-to-end testing framework with a BDD-style syntax where tests are written as user-perspective scenarios using an…
984241active
megadose/toutatis
Toutatis is a Python CLI tool that extracts public information from Instagram accounts, such as emails, phone numbers, follower counts, and…
324240active
praw-dev/praw
PRAW is a Python package that wraps Reddit's API, providing simple, rule-compliant access to Reddit data and actions. It handles OAuth auth…
984234stable
morpheus65535/bazarr
Bazarr is a self-hosted companion application to Sonarr and Radarr that automatically manages and downloads subtitles for your TV series an…
924234active
watsonbox/exportify
Exportify is a browser-based application for exporting and backing up Spotify playlists to CSV files using the Spotify Web API. It runs ent…
654201active
vogler/free-games-claimer
A Node.js automation tool that periodically claims free games and DLCs on the Epic Games Store, Amazon Prime Gaming, and GOG, plus free Unr…
734197active
kanasimi/work_crawler
A multi-language downloader application that batch-downloads web novels (converting them to EPUB) and comics from a large list of Chinese, …
664193active
Patchright
Patchright is a patched, undetected fork of the Playwright browser automation framework that evades bot-detection systems like Cloudflare. …
934191active
InstaPy/InstaPy
InstaPy is a Python library built on Selenium that automates Instagram interactions such as liking, commenting, following, and unfollowing …
2718167maintenance
speedyapply/JobSpy
JobSpy is a Python library that scrapes job postings from popular job boards like LinkedIn, Indeed, Glassdoor, Google, and ZipRecruiter con…
574165active
027xiguapi/code-box
CodeBox is a browser extension for Chrome, Edge, Firefox, and 360 browsers that enhances Chinese tech blog sites like CSDN, Zhihu, Juejin, …
604140active
dotnetcore/DotnetSpider
DotnetSpider is a .NET Standard web crawling and scraping framework that is lightweight, efficient, and cross-platform. It supports distrib…
574138active
ivre/ivre
IVRE is an open-source network recon framework written in Python that collects, stores, and analyzes network intelligence from active scann…
664119active
ericciarla/trendFinder
A self-hosted Node.js application that monitors influencer posts on Twitter/X and website changes via Firecrawl, then uses LLMs (Together A…
244116active
Lucksi/Mr.Holmes
Mr.Holmes is a Python-based OSINT (open-source intelligence) CLI tool that gathers information about usernames, domains, phone numbers, and…
534112active
nghuyong/WeiboSpider
A continuously maintained Python web scraping tool for Sina Weibo built on Scrapy and the new weibo.com API. It collects user profiles, pos…
734109active
RipMeApp/ripme
RipMe is a cross-platform Java application that bulk-downloads image albums from websites like Reddit, Imgur, Twitter, Instagram, and Tumbl…
984104active
jasonxtn/Argus
Argus is a Python-based all-in-one information gathering and reconnaissance toolkit with an interactive console and modular architecture. I…
474082active
IonicaBizau/scrape-it
scrape-it is a Node.js web scraping library with a simple, declarative API for extracting data from HTML pages, built on top of tinyreq and…
934073active
ipcjs/oh-my-userscripts
A collection of Tampermonkey/Greasyfork userscripts by ipcjs that tweak and enhance websites like Bilibili, Zhihu, Bangumi, S1, and Google.…
734065active
5rahim/seanime
Seanime is an open-source self-hosted media server for anime and manga, offering a web interface and desktop app to manage local libraries,…
864062active
React Native Upgrade Helper
A web application that shows the exact file diffs between any two React Native versions to guide app upgrades. It is built on the rn-diff-p…
754062active
fake-useragent/fake-useragent
A Python library that generates realistic, up-to-date browser user-agent strings from a bundled real-world database. It supports random or …
104049active
Guyungy/damaihelper
DamaiHelper is a multi-platform ticket-grabbing automation assistant (Damai, Taopiaopiao, Binwandao) built as a Python backend with an Ant …
744042active
hafrey1/LunaTV-config
A configuration repository and Cloudflare Workers-based CORS proxy for MoonTV/LunaTV video source APIs, with daily automated API health che…
634040active
nkanaev/yarr
yarr is a web-based RSS feed aggregator written in Go, distributed as a single dependency-free binary that runs as a desktop app or self-ho…
854031active
symfony/dom-crawler
Symfony DomCrawler is a PHP component that eases DOM navigation for HTML and XML documents. It provides a Crawler class for querying and tr…
994027stable
libredirect/browser_extension
LibRedirect is a browser extension (WebExtension) that automatically redirects requests to popular sites like YouTube, Twitter, Reddit, Ins…
884024active
imsyy/DailyHotApi
DailyHotApi is a TypeScript-based API service that aggregates trending/hot list data from many Chinese platforms (Bilibili, Weibo, Zhihu, D…
614021active

← prev page 3 / 20 next →