Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
ssili126/tv
A Python tool that automatically collects IPv4 hotel IPTV live stream sources (including CCTV, satellite, and some local Chinese channels),…
701956active
vmoranv/jshookmcp
An MCP server exposing 600+ tools across 34 domains for JavaScript reverse engineering and security research, including browser automation,…
591954active
AWeirdDev/flights
fast-flights is a Python library that scrapes Google Flights by generating Base64-encoded Protobuf query strings, returning strongly-typed …
861939active
feder-cr/invisible_playwright
A Python library that provides an antidetect, stealth-patched Firefox build for Playwright, with fingerprints set at the C++ engine level a…
811938active
linbailo/zyqinglong
A collection of self-use scripts for the Qinglong panel that automates daily check-ins and coupon collection for Chinese apps like Didi, Me…
401925active
damoeb/rss-proxy
RSS-proxy is a self-hostable web service that generates RSS, ATOM, or JSON feeds from almost any static website by analyzing its HTML struc…
321924active
SleepingBag945/dddd
dddd is a Go-based batch information gathering and supply-chain vulnerability detection CLI tool designed to streamline red team workflows.…
191924active
MarginaliaSearch/MarginaliaSearch
Marginalia Search is an independent, open-source internet search engine that indexes text-oriented, non-commercial, small and old websites.…
651923active
ZeroPointSix/outlookEmailPlus
OutlookMail Plus is a self-hosted email manager purpose-built for account registration and verification workflows. It automates fetching ve…
811917active
KEV0143/Parser-Chitai-Gorod
A Python-based scraper for the Russian online bookstore Chitai-Gorod that collects book URLs across catalog pages and extracts structured p…
291914active
aeonfun/opendia
OpenDia is an open-source browser extension (npm package) that connects your Chrome, Arc, or Firefox browser to AI models via the Model Con…
831913active
lanyeeee/bilibili-video-downloader
A cross-platform GUI desktop application built with Tauri for downloading videos, audio, subtitles, danmaku, and covers from Bilibili. It s…
751912active
extractus/article-extractor
A TypeScript library that extracts the main article content, title, image, and metadata from a given URL or raw HTML string. It supports cu…
981909active
sjdonado/idonthavespotify
A web app that converts music links between streaming services like Spotify, Apple Music, YouTube Music, Tidal, and Deezer. It parses the s…
751908active
karlicoss/promnesia
Promnesia is a browser extension plus Python backend that enhances browsing history by annotating visited pages with context from multiple …
821895active
bellingcat/octosuite
Octosuite is a terminal-based toolkit for analyzing GitHub data, usable as an interactive TUI, a CLI, or a Python library. It queries user,…
761895active
trevorhobenshield/twitter-api-client
A Python library implementing X/Twitter's v1, v2, and GraphQL APIs for automation and scraping. It supports account actions like tweeting, …
301894active
xtekky/TikTok-ViewBot
A Python-based TikTok view bot that sends fake views via HTTP requests against Zefoy, without needing Selenium. It includes an automatic ca…
351889active
nsonaniya2010/SubDomainizer
SubDomainizer is a Python CLI tool that discovers hidden subdomains and secrets in webpages, external JavaScript files, GitHub, and local f…
661886active
zema1/watchvuln
WatchVuln is a self-hosted Go service that scrapes high-quality vulnerability sources (Aliyun AVD, Chaitin, OSCS, Qianxin TI, Seebug, CISA …
621885active
scholarly-python-package/scholarly
scholarly is a Python library for retrieving author and publication metadata from Google Scholar through a friendly, Pythonic API. It handl…
561879active
degoog-org/degoog
Degoog is a self-hosted search aggregator that queries multiple search engines in parallel from your own server and merges results into a s…
801878active
404-novel-project/novel-downloader
An extensible userscript (Tampermonkey/Greasemonkey/Violentmonkey) that downloads novels from many Chinese web novel sites and exports them…
761877active
coder-hxl/x-crawl
x-crawl is a flexible Node.js crawler library that supports crawling dynamic pages, static pages, API data, and files, with optional AI ass…
661877active
joyce677/TrendRadar
TrendRadar is a lightweight self-hosted hot-topic monitor that aggregates trending news from 35 Chinese platforms (Weibo, Zhihu, Bilibili, …
621877active
null2264/yokai
Yōkai is a free and open source manga reader app for Android, forked from Tachiyomi/Mihon. It supports local and online manga reading with …
931875active
ThePhaseless/Byparr
Byparr is a self-hosted Python service that solves antibot browser challenges (like Cloudflare checks) and returns valid clearance cookies …
901865active
0xKayala/NucleiFuzzer
NucleiFuzzer is a Python-based automation tool that combines URL discovery tools (ParamSpider, Waybackurls, Gauplus, Hakrawler, Katana) wit…
661862active
tsingyuai/growth-lab
Growth Lab is an open-source, end-to-end AI growth system that runs product marketing and user-acquisition loops through natural-language c…
561861active
pmh1314520/WebRPA
WebRPA is an open-source, no-code visual RPA tool for building automation workflows by dragging and connecting modules, covering web scrapi…
821856active
ARC-MX/sgcc_electricity_new
A Home Assistant integration (deployed via Docker) that scrapes China State Grid (SGCC) accounts to fetch electricity billing and usage dat…
891855active
sarperavci/GoogleRecaptchaBypass
A Python library that automatically solves Google reCAPTCHA v2 challenges in under five seconds using browser automation with DrissionPage …
681855active
download-directory/download-directory.github.io
A web app that lets users download a single subdirectory from a GitHub repository as a zip file, filling a gap in GitHub's native functiona…
691852active
tiantianGPU/reg-factory
A Python-based local web console that automates bulk registration of email and AI service accounts (Outlook, Gmail, ChatGPT, Grok, Claude, …
801850active
ZianTT/BHYG
BHYG is a script/tool for automatically grabbing tickets for Bilibili World (BW) events via Bilibili's member purchase (会员购) platform. It a…
731849active
GuDaStudio/GrokSearch
GrokSearch is an MCP server built on FastMCP that gives Claude Code and other LLM clients real-time web access via a dual-engine architectu…
481849active
1234567Yang/cf-proxy-ex
A Cloudflare Workers-based super proxy that lets users access websites like GitHub, DuckDuckGo, and StackOverflow through a different URL w…
691848active
wapiti-scanner/wapiti
Wapiti is an open-source black-box web vulnerability scanner written in Python that crawls deployed web applications and fuzzes scripts and…
981846active
mwmbl/mwmbl
Mwmbl is an open source, non-profit web search engine with no ads or tracking, where the community determines rankings and runs distributed…
771844active
AnySearch
AnySearch is a unified real-time search engine service for AI agents, distributed as an agent skill package and an MCP server. It provides …
581838active
abinthomasonline/repo2txt
A browser-based tool that converts GitHub, GitLab, Azure DevOps repositories, local folders, or ZIP files into a single formatted text file…
651835active
microlinkhq/browserless
A Node.js library that wraps Puppeteer to provide a production-ready headless Chrome/Chromium driver with built-in screenshot, PDF generati…
951831active
78778443/QingScan
QingScan is a self-hosted, open-source security operations platform that unifies vulnerability scanning, code auditing, asset inventory, an…
661831active
tryolabs/requestium
Requestium is a Python library that merges Requests, Selenium, and Parsel into a single integrated tool for web automation. It lets scripts…
771830active
lds133/weather_landscape
A Python application that renders weather forecasts as a stylized landscape image instead of numeric dashboards, encoding time, temperature…
561827active
initstring/linkedin2username
A Python OSINT tool that scrapes LinkedIn employee lists for a target company and generates multiple probable username formats (e.g., first…
761825active
1N3/BlackWidow
BlackWidow is a Python-based web application spider that crawls a target site to collect URLs, dynamic parameters, subdomains, email addres…
571821active
autoclaw-cc/xiaohongshu-skills
A set of AI agent skills (SKILL.md format) plus a Chrome extension that automates Xiaohongshu (RED) using your real logged-in browser sessi…
701819active
petronny/gfwlist2pac
A tool that automatically converts the gfwlist proxy rules into a PAC (Proxy Auto-Config) file every day. The generated gfwlist.pac is serv…
771816active
enetx/surf
Surf is an advanced HTTP client library for Go with fluent, chainable API design. It supports browser impersonation (Chrome/Firefox), JA3/J…
831808active
POf-L/Fanqie-novel-Downloader
A cross-platform desktop and mobile application built with Rust and Tauri v2 for searching, reading, and downloading novels from Fanqie Nov…
821801active
LagradOst/QuickNovel
QuickNovel is a free, ad-free, open-source Android app for downloading novels from many web novel sites, which also functions as an EPUB re…
991786active
emacs-elfeed/elfeed
Elfeed is an extensible web feeds client for Emacs supporting Atom, RSS, and JSON Feed formats. It provides a search-based UI inspired by n…
671784active
kost/dvcs-ripper
dvcs-ripper is a set of Perl command-line tools that download (rip) web-accessible version control repositories such as GIT, SVN, Mercurial…
321784stable
Masterminds/html5-php
A standards-compliant HTML5 parser and serializer written entirely in PHP. It parses HTML5 documents and fragments into standard PHP DOM ob…
871782stable
mdc-ng/mdc-ng
A self-hosted media metadata scraper and organizer for adult video libraries, written in Rust with a Next.js web UI. It scrapes metadata fr…
831782active
ShunCai/QZoneExport
A browser extension (Chrome/Edge/Firefox, Manifest V3) that backs up QQ Zone data—posts, blogs, private diaries, albums, videos, comments, …
881780active
ScrapeCreators/social-media-research-skills
A collection of AI agent skills for social media research built on the ScrapeCreators scraping API. The skills give agents complete workflo…
581777active
bulianglin/psub
psub is a proxy subscription conversion tool deployed on Cloudflare Workers that acts as a reverse proxy for subscription conversion backen…
101772active
inulute/medium-unlocker
Medium Unlocker is an Android app (with an accompanying web frontend) that bypasses Medium's paywall by redirecting article URLs to the fre…
911767active
deweizhu/bookget
bookget is a Go-based command-line tool for downloading digitized ancient books and rare texts from 50+ digital libraries. It ships prebuil…
601764active
cambecc/earth
Earth is a JavaScript web application that visualizes global weather, wind, ocean currents, and related atmospheric conditions on an animat…
326588maintenance
LoseNine/ruyipage
RuyiPage is a Python browser automation framework built on Firefox and the WebDriver BiDi protocol, shipping with an anti-detection Firefox…
791759active
josh0xA/darkdump
Darkdump is an open-source OSINT tool for querying multiple dark web search engines and scraping onion site results for emails, metadata, k…
701757active
website-scraper/node-website-scraper
A Node.js library that downloads entire websites to a local directory, including HTML, CSS, images, and JavaScript assets. It parses HTTP r…
721751active
adsbypasser/adsbypasser
AdsBypasser is a lightweight userscript that automatically skips countdown ads, continue/redirect pages, and prevents ad pop-up windows acr…
991750active
Aas-ee/open-webSearch
Open-WebSearch is a TypeScript tool providing an MCP server, CLI, and local daemon for multi-engine web search and content retrieval withou…
791747active
sindresorhus/pageres-cli
A Node.js command-line tool that captures screenshots of websites at multiple resolutions using headless Chrome (Puppeteer). It is useful f…
441743active
IvanGlinkin/Fast-Google-Dorks-Scan
A shell-based OSINT tool that automates Google dork searches against a target website to uncover admin panels, exposed file types, and path…
461742active
egoist/sitefetch
A Node.js CLI tool that crawls an entire website and saves its pages as a single text file, using Mozilla Readability to extract clean cont…
221736active
dilame/instagram-private-api
A NodeJS/TypeScript SDK providing a client for Instagram's private (undocumented) API, enabling full programmatic access to feeds, direct m…
236469maintenance
cxOrz/chaoxing-signin
A Node.js/TypeScript tool that automates sign-in for the Chaoxing (Superstar Learning) online course platform, supporting normal, photo, ge…
101724active
hzm0321/real-time-fund
A Next.js web application for real-time mutual fund valuation and top-holdings stock tracking, primarily for Chinese funds, with portfolio,…
781723active
claffin/cloudproxy
CloudProxy is a self-hosted Python tool that provisions and manages proxy servers across multiple cloud providers, rotating IPs to improve …
811722active
Jules-WinnfieldX/CyberDropDownloader
A Python-based bulk downloader that scrapes and downloads files from Cyberdrop.me and dozens of other file hosts and image galleries. It is…
101720active
jldbc/pybaseball
pybaseball is a Python package that scrapes and retrieves current and historical baseball statistics from sources like MLB Statcast (Baseba…
501719active
qinlili23333/ctfileGet
A web-based resolver that obtains one-time direct download URLs for files hosted on Chengtong Network Disk (ctfile/城通网盘) using its official…
551714active
hangone/WeBan
WeBan is a Python-based automation tool that automatically completes courses and exams on the Weiban (安全微伴) university safety education pla…
851712active
3441293738/creatorhub
CreatorHub is a self-hosted web panel built with Python and FastAPI for managing, monitoring, scraping, downloading, and publishing content…
581712active
zstmfhy/zlibrary-to-notebooklm
A Python CLI tool that automatically downloads books from Z-Library and uploads them to Google NotebookLM in one command. It uses Playwrigh…
431711active
Python3Spiders/WeiboSuperSpider
A Weibo (Chinese microblog) scraping toolbox in Python covering users, topics, and comments, with extras like image downloading, sentiment …
751705active
aisingapore/TagUI
TagUI is a free, open-source robotic process automation (RPA) tool from AI Singapore that lets users write simple text flows to automate re…
656325maintenance
utkusen/urlhunter
urlhunter is a Go-based recon CLI tool that searches URLs exposed via shortener services like bit.ly and goo.gl. It downloads daily URLTeam…
241697active
oxylabs/google-play-scraper
A free Python-based Google Play Store scraper that collects public app, movie, and book data via search queries. It is a companion tool to …
671693active
openkursar/hello-halo
Halo is an open-source desktop AI workstation that wraps frontier coding agents like Claude Code and Codex in a visual GUI, with a pluggabl…
811688active
amaancoderx/npxskillui
SkillUI is a Node.js CLI that crawls websites, git repos, or local codebases and extracts their complete design system (colors, typography,…
511688active
LearnPrompt/ai-news-radar
AI News Radar is an automated 24-hour AI/tech news aggregator that fetches sources, deduplicates stories, and scores headlines with three p…
801687active
Ge0rg3/requests-ip-rotator
A Python library that mounts AWS API Gateway as a proxy onto requests sessions, rotating source IPs on every request using AWS's large IP p…
901675active
jarun/googler
googler is a Python command-line tool for performing Google web, news, and video searches directly from the terminal. It displays titles, U…
106204maintenance
vibheksoni/stealth-browser-mcp
A Python MCP server that exposes stealth browser automation (via nodriver and Chrome DevTools Protocol) to AI agents, letting them navigate…
611674active
rushter/selectolax
Selectolax is a fast Python HTML5 parser library written in Cython, binding to the Modest and Lexbor parsing engines. It provides CSS selec…
941665active
NeteaseCloudMusicApiEnhanced/api-enhanced
A Node.js API service providing comprehensive access to NetEase Cloud Music (网易云音乐) endpoints, a half-refactored and enhanced fork of the p…
841661active
ReScienceLab/opc-skills
A collection of open-source Agent Skills (folders of instructions, scripts, and resources) designed for solopreneurs, indie hackers, and on…
731658active
itteco/iframely
Iframely is a self-hosted oEmbed proxy and URL metadata service that takes any URL and returns rich media embed codes and page metadata. It…
671649active
nasa-gibs/worldview
NASA Worldview is an interactive web application for browsing over 1000 global, full-resolution satellite imagery layers from NASA's Global…
951648active
srx-2000/spider_collection
A collection of Python web crawler scripts targeting sites like Bilibili, Zhihu, Weibo, NetEase Music, GitHub, and Anjuke, built with reque…
321646active
Benexl/yt-x
A POSIX-compliant shell script that lets you browse YouTube and other yt-dlp-supported sites from the terminal using fzf or from an app lau…
851645active
itsToggle/plex_debrid
A Python application that monitors Plex, Trakt, and Overseerr watchlists and automatically fetches requested movies and shows from torrent …
101644active
Jesseovo/last30days-skill-cn
An AI Agent skill (for Claude Code / OpenClaw) that automatically searches content from the last 30 days across 8 major Chinese internet pl…
561642active
Sparticuz/chromium
A TypeScript library that packages a serverless-optimized Chromium binary (Brotli-compressed) with decompression code and predefined launch…
951641active

← prev page 7 / 20 next →