Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
EtherDream/jsproxy
An online web proxy that uses browser Service Workers to intercept and rewrite requests client-side, so the nginx-based server only forward…
329327maintenance
apify/agent-skills
A collection of production-grade agent skills from Apify that give AI coding agents (Claude Code, Cursor, Windsurf, Codex, Gemini CLI) expe…
592361active
tavily-ai/tavily-mcp
A Model Context Protocol (MCP) server that exposes Tavily's web search, page extraction, site mapping, and crawling tools to AI assistants …
652354active
vitiko98/qobuz-dl
A Python command-line tool for searching, exploring, and downloading lossless and Hi-Res music (FLAC/MP3) from Qobuz. It supports interacti…
372351active
lcomplete/huntly
Huntly is a self-hosted, AI-powered personal information hub that captures, archives, and organizes web content, RSS feeds, tweets, and hig…
862343active
geekgeekrun/geekgeekrun
GeekGeekRun is a free, open-source desktop application (built on Electron, Vue, Puppeteer, SQLite/TypeORM) that automates job hunting on th…
792341active
FriendsOfPHP/Goutte
Goutte is a PHP screen scraping and web crawling library providing a simple API to crawl websites and extract data from HTML/XML responses.…
109192maintenance
miraclx/freyr-js
Freyr is a Node.js CLI tool that downloads songs from music streaming services like Spotify, Apple Music, and Deezer. It extracts track met…
672330active
erengy/taiga
Taiga is an open-source Windows desktop application that automatically detects anime videos you watch on your computer and syncs your progr…
672329active
gxr404/yuque-dl
yuque-dl is a Node.js CLI tool that downloads Yuque (语雀) knowledge bases and documents as local Markdown files. It supports batch downloads…
702327active
EricZhu-42/SteamTradingSiteTracker
A Steam skin trading market tracker that continuously monitors sell-to-cash (挂刀) ratios across BUFF, IGXE, C5, UUYP, and ECO platforms for …
322324active
easychen/checkchan-dist
Check酱 (CheckChan) is a web page content change monitoring tool consisting of a Chrome/Edge browser extension and a self-hostable cloud com…
322321active
dataabc/weibo-search
A Python/Scrapy-based crawler that continuously fetches Weibo keyword and hashtag search results, including full post metadata, images, and…
712317active
chaolucky18/xuexitongScript
A Tampermonkey userscript (also runnable via browser console) that automates course video playback on the Xuexitong (Chaoxing) learning pla…
742311active
sjdirect/abot
Abot is an open source C# web crawler framework built for speed and flexibility, handling multithreading, HTTP requests, scheduling, and li…
742310active
jaywcjlove/github-rank
A GitHub user and repository ranking site that publishes global and China leaderboards of GitHub users by followers and repositories by sta…
982308active
0xMassi/webclaw
webclaw is a Rust-based web extraction toolkit that turns any URL into clean, LLM-ready markdown, JSON, or token-optimized text, including …
772305active
NAStool/nas-tools
A self-hosted NAS media library management tool that automates organizing, scraping, and syncing media content for media servers like Plex,…
109027maintenance
Anakin-Inc/anakin
AnakinScraper OSS is a self-hosted web scraping API written in Go that turns any website into LLM-ready markdown or structured JSON via a s…
732302active
Xiangyu-CAS/xiaohongshu-ops-skill
A skill for the OpenClaw agent that turns it into a Xiaohongshu (RedNote) operations assistant, using browser automation (CDP) to analyze f…
502291active
anaskhan96/soup
soup is a small Go library for web scraping with an API modeled after Python's BeautifulSoup. It fetches HTML over HTTP and builds a DOM th…
962286active
zenbu-labs/terminal-browser
terminal-browser is a real Chromium-based browser that renders inside your terminal using the kitty graphics protocol and Electron's offscr…
802285active
Future-Scholars/paperlib
Paperlib is an open-source, cross-platform desktop application for managing academic papers, built with TypeScript and Electron. It scrapes…
602274active
saifyxpro/HeadlessX
HeadlessX is a self-hosted browser automation and web scraping platform powered by Camoufox (a C++-patched Firefox) to bypass anti-bot syst…
762268active
ajayyy/DeArrow
DeArrow is an open-source browser extension that crowdsources better titles and thumbnails for YouTube videos, replacing sensationalized cl…
962259active
Johnserf-Seed/TikTokDownload
A Python CLI tool for batch-downloading Douyin (Chinese TikTok) content without watermarks, including user profile posts, likes, favorites,…
238815maintenance
ying-ck/fanqienovel-downloader
A Python tool that downloads novels from Fanqie Novel (fanqienovel.com) by URL or book ID, with search, batch download, update, and backup …
892256active
MinhasKamal/DownGit
DownGit is a web tool that creates direct download links for any public GitHub directory or file, packaging them as a zip archive. It is ho…
542256active
coleam00/mcp-crawl4ai-rag
An MCP server that combines Crawl4AI web crawling with RAG capabilities backed by a Supabase vector database, exposing tools for AI agents …
342245active
vasu-devs/JustHireMe
JustHireMe is a local-first desktop workbench (Tauri frontend, Python backend) that scrapes job postings, ranks role fit against your profi…
782236stable
goclone-dev/goclone
Goclone is a Go CLI utility that downloads entire websites to a local directory, preserving relative link structure so the mirrored site ca…
622231active
dou-jiang/codex-console
A Python-based integrated console for automating OpenAI/Codex account registration, login, token retrieval, subscription management, and up…
722230active
wabarc/wayback
Wayback is an open-source web archiving tool written in Go that captures and preserves web pages via services like Internet Archive, archiv…
852227active
hhursev/recipe-scrapers
A Python library for extracting structured recipe data (title, ingredients, instructions, cooking times, images, nutrients) from cooking we…
982216active
VermiIIi0n/fuckZHS
A Python 3 automation script that automatically completes Zhihuishu (智慧树) online course videos, including auto-answering pop-up quiz questi…
452216active
Hubs-Foundation/hubs
Hubs is an open-source, browser-based multi-user 3D virtual world and social VR platform built with A-Frame, Three.js, and WebXR/WebRTC. It…
822214active
AnotiaWang/deep-research-web-ui
A web UI for the dzhng/deep-research project that performs iterative, AI-driven research by combining search engines, web scraping, and LLM…
662205active
ReaJason/xhs
A Python SDK that wraps requests to the Xiaohongshu (Little Red Book) web platform for extracting data. It provides a programmatic client f…
432202active
Imangazaliev/DiDOM
DiDOM is a fast and simple PHP library for parsing and manipulating HTML and XML documents. It supports loading from strings, files, or URL…
522198active
mampfes/hacs_waste_collection_schedule
A Home Assistant custom integration that fetches waste collection schedules from many service providers, ICS/iCal files, or user-defined da…
982195active
opennaslab/kubespider
Kubespider is a self-hosted download orchestration system that turns an idle Linux server into a NAS download center. It uses pluggable sou…
572190active
Owez/yark
Yark is a Python CLI tool for archiving YouTube channels, downloading videos and accumulating metadata over time with change reports. It in…
662184active
levigross/grequests
GRequests is a Go library that wraps net/http with a convenient, Python Requests-style API. It provides helpers for all HTTP verbs, JSON/XM…
732182active
bigintpro/csdn_downloader
A Java (Spring Boot + Dubbo) based web service that downloads CSDN resources such as articles and paid/VIP documents without needing points…
592179active
NobyDa/Script
A collection of JavaScript scripts and configuration files for iOS proxy tools such as Surge, Quantumult X, Loon, Stash, and Shadowrocket. …
768448maintenance
ericchiang/pup
pup is a command line tool for parsing and filtering HTML using CSS selectors, inspired by jq. It reads HTML from stdin, applies selector-b…
238435maintenance
zhzyker/dismap
Dismap is a Go-based asset discovery and identification tool that fingerprints web, TCP, UDP, and TLS services using a rule base of 4500+ w…
232163active
AAndyProgram/SCrawler
SCrawler is a Windows GUI application that downloads photos and videos from user profiles across many social media and content sites, inclu…
932154active
philss/floki
Floki is an Elixir HTML parser that lets you search document nodes using CSS selectors. It supports multiple parsing backends (mochiweb_htm…
862149stable
AaronL725/grok-register
A Python toolkit that automates bulk registration of Grok accounts using real Chromium browser automation, with GUI, CLI, and WebUI interfa…
582149active
bit4woo/domain_hunter_pro
Domain Hunter Pro is a Burp Suite plugin (Java jar) for automated domain and subdomain collection, web title fetching, and target managemen…
632145active
Rongronggg9/RSS-to-Telegram-Bot
A self-hosted Telegram bot that delivers RSS/Atom feed updates to Telegram chats with rich-text formatting and media support. It is multi-u…
672141active
up209d/ResourcesSaverExt
A Chrome browser extension that downloads all resources of a website with one click while preserving the original folder structure. It inte…
372141active
iuroc/bilidown
Bilidown is a Bilibili video parsing and downloading desktop application supporting 8K video, Hi-Res audio, Dolby Vision, batch parsing, QR…
832135active
pablouser1/ProxiTok
ProxiTok is an open-source alternative frontend for TikTok, inspired by Nitter, written in PHP. It proxies all requests to TikTok server-si…
342133active
mgz0227/legado-Harmony
Legado (开源阅读) for HarmonyOS is a free, open-source novel and ebook reader application. It supports custom book sources with user-defined sc…
922125active
benvinegar/counterscale
Counterscale is a self-hosted, privacy-friendly web analytics tracker and dashboard built on Cloudflare Workers and Workers Analytics Engin…
792122active
cxfksword/jellyfin-plugin-metashark
A Jellyfin metadata plugin that scrapes movie and anime metadata primarily from Douban, with TheMovieDb used to fill in missing episode dat…
972119active
feedjira/feedjira
Feedjira is a Ruby library for parsing syndication feeds such as RSS and Atom. It supports extensible and custom parsers, letting users add…
772103active
drunkdream/weread-exporter
A Python CLI tool that exports books from WeChat Read (微信读书) into epub, pdf, and mobi formats. It hooks into the web reader's Canvas render…
642101active
yuanzl77/IPTV
A Python tool that aggregates IPTV live streaming sources daily, validates them with HTTP checks and FFprobe quality probing, and generates…
692100active
chao325/MaoTai_GUIT
A Windows GUI/console application for automated flash-sale purchasing (sniping) on JD, Taobao, and Damai, distributed as a ready-to-run EXE…
722097active
vvoovv/blosm
Blosm is a Blender addon that imports real-world geodata—OpenStreetMap buildings and roads, Google 3D city tiles, and terrain—with global c…
662097active
ipfs/public-gateway-checker
A web application that displays a list of public IPFS gateways and checks whether each is online, including CORS, IPNS, origin isolation, a…
982096active
zorlan/skycaiji
SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a …
772089active
hanc00l/nemo_go
Nemo is an automated information-gathering platform for penetration testing that integrates common recon tools (Masscan, Nmap, Subfinder, H…
882086active
prajwalch/TorrentSearch
TorrentSearch is an Android app that searches torrents across multiple providers simultaneously, with category filters, detailed results, b…
862084active
AzizKpln/Moriarty-Project
Moriarty Project is a web-based phone number investigation tool written in Python that gathers information about a given phone number. It a…
232074active
WebReflection/linkedom
LinkeDOM is a triple-linked-list based DOM implementation for DOM-less environments like Node.js and Deno, closely following the DOM standa…
742070active
towfiqi/serpbear
SerpBear is an open-source, self-hosted search engine position tracking app for monitoring website keyword rankings in Google. It scrapes S…
832064active
lexbor/lexbor
Lexbor is a fast, standards-compliant HTML parser and DOM library written in pure C99, with additional modules for CSS, URL, encoding, and …
822055active
oxylabs/how-to-scrape-google-images
A Python-based command-line tool that scrapes Google Images search results, including reverse image search based on a provided image URL. I…
642055active
upbit/pixivpy
PixivPy is a Python client library for the Pixiv App-API, supporting authenticated access via refresh tokens. It provides methods for searc…
382052active
oxylabs/how-to-scrape-google-flights
A Python-based free scraper tool and tutorial for extracting flight data (prices, times, airlines) from Google Flights pages, either direct…
582048active
probberechts/soccerdata
A Python library of scrapers that collect soccer data from popular websites like FBref, ESPN, WhoScored, Sofascore, SoFIFA, Understat, Club…
932040active
wbt5/real-url
A Python collection of scripts that extracts real streaming URLs (live stream sources) and danmaku (bullet comments) from 59 Chinese and in…
327827maintenance
rubycdp/ferrum
Ferrum is a Ruby library providing a clean, high-level API to control Chrome or Chromium via the Chrome DevTools Protocol (CDP), with no Se…
902037active
pt-plugins/PT-Plugin-Plus
PT-Plugin-Plus is a Web Extensions browser plugin for Chrome, Edge, and Firefox that streamlines using private tracker (PT) sites, enabling…
107814maintenance
oxylabs/how-to-scrape-amazon-prices
A Python-based example repository and free CLI tool for scraping Amazon product prices, best sellers, search results, and deals from depart…
642028active
Ocyss/boss-helper
A browser extension (also available as a userscript) that enhances the Boss Zhipin job platform by removing ads, improving the UI, enabling…
912027active
TrianguloY/URLCheck
URLCheck is an open-source Android app that acts as an intermediary when opening URLs, letting users inspect, clean, and modify links befor…
872027active
elliotgao2/gain
Gain is an asynchronous web crawling framework for Python built on asyncio, aiohttp, and lxml/pyquery. Users declare items and parsers decl…
752019active
afar1/fieldtheory-cli
A TypeScript CLI that syncs X/Twitter bookmarks to local markdown files, provides BM25 full-text search and LLM-based classification, and e…
582015active
john-kurkowski/tldextract
A Python library that accurately splits URLs into subdomain, domain, and public suffix components using the Public Suffix List. It also shi…
862014stable
jonhoo/fantoccini
Fantoccini is a Rust library providing a high-level async API for programmatically controlling browsers via the WebDriver protocol. It supp…
752014active
dmzz-yyhyy/LightNovelReader
LightNovelReader is an open-source Android light novel reader app built with Kotlin and Jetpack Compose, featuring a lightweight footprint …
892012active
jimmc414/onefilellm
OneFileLLM is a Python command-line tool and library that aggregates content from sources like GitHub repos, pull requests, arXiv/Sci-Hub p…
692011active
watercrawl/WaterCrawl
WaterCrawl is a self-hostable web application (Python/Django/Scrapy/Celery) that crawls websites and transforms web content into LLM-ready …
822010active
AnswerOverflow/AnswerOverflow
Answer Overflow is an open-source application that indexes Discord server threads into searchable, SEO-friendly web pages so community know…
652007active
supermemoryai/markdowner
Markdowner is a fast web service that converts any website into LLM-ready markdown, with optional LLM filtering, detailed responses, and au…
241999active
nottelabs/notte
Notte is a full-stack framework and cloud platform for building, deploying, and scaling AI web agents and browser automations. It combines …
841997active
ChinaGodMan/UserScripts
A collection of Tampermonkey/Greasyfork userscripts modified from the internet, written in JavaScript. The scripts add browser enhancements…
681997active
zhegexiaohuozi/SeimiCrawler
SeimiCrawler is an agile, standalone, distributed Java crawler framework inspired by Python's Scrapy, with deep Spring Boot integration and…
851990active
karpathy/jobs
A research tool that scrapes the Bureau of Labor Statistics Occupational Outlook Handbook (342 occupations) and renders an interactive tree…
471989active
A9T9/RPA
Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible …
961985active
danmactough/node-feedparser
A Node.js library for parsing RSS, Atom, and RDF syndication feeds as a streaming interface. It resolves relative URLs and correctly handle…
701976stable
ViennaRSS/vienna-rss
Vienna is a free, open-source RSS/Atom/JSON feed newsreader for macOS with a native Apple Mail-like interface. It can fetch feeds directly …
971972active
walkingddd/TgtoDrive
TgtoDrive is a self-hosted, Docker-deployed media automation platform that chains resource discovery (Telegram channel monitoring, search s…
751964active
Flexget/Flexget
FlexGet is a multipurpose Python automation tool for content like torrents, NZBs, podcasts, comics, series, and movies. It pulls from sourc…
951963active
Tabula
Tabula is a local web application for extracting data tables from text-based PDF files into CSV, Excel, or JSON. It is powered by the tabul…
287472maintenance

← prev page 6 / 20 next →