Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
blacklanternsecurity/MANSPIDER
MANSPIDER is a Python CLI tool that crawls SMB shares across entire networks to find files by filename or content, with regex support and t…
771406active
superhedgy/AttackSurfaceMapper
AttackSurfaceMapper is a Python CLI reconnaissance tool that expands a target's attack surface using OSINT and active techniques like subdo…
321405active
nianzhibai/91
A self-hosted private video site written in Go that aggregates videos from multiple cloud drives (115, PikPak, 123pan, OneDrive, Google Dri…
801404active
KoalaBear84/OpenDirectoryDownloader
A cross-platform C#/.NET command-line tool that indexes open directory listings across 130+ supported formats, including FTP(S), Google Dri…
941389active
lorey/mlscraper
mlscraper is a Python library that automatically extracts structured data from HTML pages using machine learning. Instead of writing CSS se…
231385active
Avnsx/fansly-downloader
A Python-based tool for bulk downloading photos, videos, and audio from fansly.com, also shipped as a standalone Windows executable. It sup…
101384active
thinh-vu/vnstock
Vnstock is an open-source Python library for extracting and analyzing Vietnam stock market data, returning data as pandas DataFrames via si…
901383active
jachinlin/geektime_dl
A Python CLI tool that downloads Geektime (极客时间) courses and converts them into ebooks for reading on Kindle. It handles login, course quer…
661381active
CIRCL/AIL-framework
AIL framework is an open-source Python platform for collecting, crawling, processing, and analyzing unstructured data from the clear web, T…
671378active
mattsse/chromiumoxide
chromiumoxide is a Rust library providing a high-level async API for controlling Chrome or Chromium via the Chrome DevTools Protocol. It ca…
751375active
okfn-brasil/querido-diario
Querido Diário is an open-source project by Open Knowledge Brasil that scrapes and aggregates Brazilian municipal official gazettes (diário…
751373active
tavily-ai/tavily-python
The official Python SDK for the Tavily API, providing search, content extraction, crawling, site mapping, and research capabilities. It let…
721373active
Minerchu/dongguaTV
A self-hosted Node.js video aggregation platform that searches 30+ movie/TV resource site APIs, aggregates results, and scrapes metadata fr…
401372active
xjbeta/iina-plus
IINA+ is a small macOS application that adds danmaku (bullet comment) support and live-stream/video playback for Chinese platforms to the I…
971370active
dethcrypto/dethcode
DethCode lets you view the verified source of deployed Ethereum smart contracts in an ephemeral VS Code instance by changing an Etherscan U…
521370active
fasnow/fine
Fine is a Chinese-language cyberspace asset mapping and reconnaissance tool integrating FOFA, Hunter, Quake, ZoomEye, and Shodan APIs, plus…
821366active
python273/vk_api
vk_api is a Python library that wraps the VKontakte (vk.com) API, simplifying authentication and API method calls for building scripts and …
831364active
xlang-ai/OpenAgents
OpenAgents is an open platform for using and hosting LLM-powered language agents, featuring a Data Agent for Python/SQL analysis, a Plugins…
284859maintenance
Maasea/sgmodule
A collection of Surge modules (sgmodule files) written in JavaScript that enhance or modify mobile apps and network behavior, such as remov…
741359active
xiaohucode/xiangse
A curated collection of video and manga source plugins (.xbs files) for the Xiangse Guige (香色闺阁) reading/media app, imported via URL. It ag…
101359active
SplashtopInc/winstall
winstall is a free, open-source web app for browsing and searching Microsoft's Windows Package Manager (winget) repository of 14,600+ apps.…
771358active
firecrawl/open-scouts
Open Scouts is an AI-powered web monitoring platform where users create automated 'scouts' that run on a schedule to search the web and sen…
531358active
fugary/calibre-douban
A Calibre metadata source plugin that fetches book metadata from Douban by crawling book.douban.com web pages, since Douban no longer offer…
631357active
scrapy/parsel
Parsel is a BSD-licensed Python library for extracting data from HTML, XML, and JSON documents using CSS selectors, XPath expressions, JMES…
881352active
raznem/parsera
Parsera is a lightweight Python library for scraping websites using LLMs, letting users define elements to extract with natural-language de…
541350active
nmdias/FeedKit
FeedKit is a Swift library for parsing and generating RSS, Atom, and JSON Feed formats. It supports common namespaces like Dublin Core, Med…
991348active
MiniGlome/Archive.org-Downloader
A Python 3 command-line script that downloads borrowable books from archive.org and Open Library and assembles them into PDF files. It requ…
761348active
LeetaoGoooo/RSSAid
RSSAid is a Flutter-based mobile app that complements RSSHub by helping users discover and subscribe to RSS feeds from websites, similar to…
761345active
philippta/flyscrape
Flyscrape is a standalone command-line web scraping tool written in Go that lets users write extraction logic in JavaScript with a jQuery-l…
421345active
SpiderClub/weibospider
A distributed web crawler for Sina Weibo (Chinese microblogging platform) built with Python, Celery, and requests. It scrapes user profiles…
324793maintenance
iota9star/mikan_flutter
A third-party Flutter mobile client for the Mikan Project (mikanani.me), an anime/seasonal bangumi torrent subscription site. It lets users…
991343active
mvdbos/php-spider
A configurable and extensible PHP web spider library for crawling websites. It supports breadth-first and depth-first traversal, URI discov…
751341active
jocmp/capyreader
Capy Reader is a free, open-source RSS feed reader and news aggregator for Android built with Kotlin and Jetpack Compose. It syncs with Fee…
891334active
DUpdateSystem/UpgradeAll
UpgradeAll is a free, open-source Android app that checks for updates to installed Android apps, Magisk modules, and other software from a …
671334active
taf2/curb
Curb provides Ruby-language bindings for libcurl, the fully-featured client-side URL transfer library, supporting both easy and multi modes…
751331active
mariostoev/finviz
An unofficial Python library for scraping FinViz.com, providing stock data, news, insider transactions, analyst price targets, and a stock …
601331active
ytdl
node-ytdl-core is a JavaScript library for downloading YouTube videos in Node.js, exposing a stream-friendly API with format selection, byt…
464730maintenance
tophubs/TopList
TopList (今日热榜) is a self-hosted aggregation website that collects trending headlines from popular sites like Zhihu, Hupu, and V2EX. It is w…
324730maintenance
sockysec/Telerecon
Telerecon is a Python-based OSINT reconnaissance framework for researching and investigating Telegram. It scrapes user profiles, messages, …
281324active
cinemagoer/cinemagoer
Cinemagoer (formerly IMDbPY) is a Python package for retrieving and managing IMDb data about movies, people, characters, and companies from…
951323active
MrTuxx/SocialPwned
SocialPwned is a Python-based OSINT tool that harvests emails published on Instagram, LinkedIn, and Twitter to find credential leaks via Pw…
101320active
monosans/proxy-scraper-checker
A fast async Rust CLI tool that scrapes HTTP, SOCKS4, and SOCKS5 proxies from arbitrary text, HTML, or JSON sources, verifies each proxy by…
771318active
lucahammer/tweetXer
A JavaScript userscript that bulk-deletes all your tweets on X (Twitter) for free, using your official Data Export file. It runs in the bro…
551316active
darwin-lau/langmanus
LangManus is a community-driven Python framework for building hierarchical multi-agent AI automation systems, where a supervisor coordinate…
251316active
Bloggify/github-calendar
A JavaScript library that embeds a user's GitHub contributions calendar into any web page with a single function call. It fetches contribut…
441315active
jmerle/competitive-companion
A browser extension that parses competitive programming problems from online judges like Codeforces, AtCoder, and Kattis. It extracts test …
871309active
codingo/VHostScan
VHostScan is a Python-based virtual host scanner that discovers hidden vhosts on a web server using wordlists, reverse lookups, and catch-a…
391309active
karust/openserp
OpenSERP is a self-hosted, MIT-licensed SERP API and CLI written in Go that returns structured search results from Google, Bing, Yandex, Ba…
911306active
yasserg/crawler4j
crawler4j is an open-source web crawler library for Java that provides a simple interface for building multi-threaded web crawlers in minut…
234618maintenance
dwisiswant0/go-dork
go-dork is a fast command-line dork scanner written in Go that automates Google dorking across multiple search engines. It supports Google,…
231301stable
DimiMikadze/orca
Orca is an AI agent application for deep LinkedIn profile analysis that scrapes posts, comments, reactions, and interaction networks, then …
751297active
bookstairs/bookhunter
bookhunter is a Go command-line tool for scraping and downloading ebooks from sources like Talebook, SoBooks, Telegram channels, and China'…
541295active
h4r5h1t/webcopilot
WebCopilot is a Bash-based automation script for bug bounty reconnaissance that enumerates subdomains using multiple tools, filters paramet…
231295active
yjl9903/AnimeGarden
AnimeGarden is a third-party mirror and aggregation site for 動漫花園 (dmhy) anime BT torrents, offering a web UI, open REST API, RSS feeds, an…
871294active
techwithtim/Price-Tracking-Web-Scraper
A full-stack price tracking application that scrapes product prices (currently Amazon.ca) using Playwright and Bright Data's Scraping Brows…
291294active
Steamauto/Steamauto
Steamauto is a free, open-source Python application that fully automates buying, selling, and delivery of CS2/CSGO skins across Steam and C…
961292active
foxhui/WebAI2API
WebAI2API is a self-hosted Node.js service that exposes web-based AI services (LMArena, Gemini, ChatGPT, DeepSeek, etc.) as OpenAI-compatib…
571292active
vega-org/vega-app
Vega is an open-source Android media streaming app built with React Native and TypeScript that lets users stream and download video content…
911290active
zohaibbashir/Google-Maps-Scrapper
A Python CLI script built on Playwright that scrapes Google Maps listings to extract business details such as name, address, website, phone…
651289active
oxylabs/paid-proxy-servers
A promotional GitHub repository for Oxylabs' commercial paid proxy services, covering residential, mobile, datacenter, ISP, and SOCKS5 prox…
591289active
P1-Team/AlliN
AlliN is a flexible, dependency-free Python scanner designed to assist penetration testing projects, especially lateral movement and intran…
401288active
misiektoja/instagram_monitor
A Python-based OSINT tool that tracks Instagram users' activities in real time, including story updates, profile changes, and follower shif…
911286active
denho/faved
Faved is a free, open-source, self-hostable bookmark manager with customizable nested tags, instant search, and duplicate detection, built …
831282active
dvcoolarun/web2pdf
A Python command-line tool that converts webpages into formatted PDFs using WeasyPrint. It supports batch conversion, recursive same-domain…
671281active
TheBeastLT/torrentio-scraper
Torrentio is a Stremio addon ecosystem that scrapes public torrent providers and serves the results as Stremio stream results. The reposito…
771280active
hadynz/obsidian-kindle-plugin
An Obsidian plugin that syncs Kindle notes and highlights into your vault, either by screen-scraping Amazon's Kindle Reader cloud library o…
751277active
sardanioss/httpcloak
httpcloak is a Go HTTP client library that reproduces browser-identical TLS, HTTP/2, and HTTP/3 fingerprints (JA3/JA4, Akamai, header order…
611276active
brunosimon/my-room-in-3d
A 3D interactive recreation of the author's room built with Three.js and JavaScript, viewable in the browser. It is a creative demo/portfol…
324486maintenance
acgotaku/YAAW-for-Chrome
A Chrome extension providing a web frontend dashboard for the Aria2 download manager via its JSON-RPC interface. It can intercept browser d…
711270active
kkangert/kspider
Kspider is a self-hosted visual web scraping platform written in Java where users define crawler workflows as flowcharts without writing ba…
141269active
umutxyp/MusicBot
Beatra is an advanced Discord music bot built on discord.js v14 that streams music from YouTube, Spotify, SoundCloud, and direct links with…
701266active
mvdan/xurls
A Go library and CLI tool that extracts URLs from arbitrary text using regular expressions built from TLD lists. It offers Relaxed and Stri…
671265active
cclank/news-aggregator-skill
A Python-based agent skill that aggregates news from 44+ sources (tech, finance, AI, international) and generates AI-summarized daily brief…
531263active
0xHJK/music-dl
A Python 3 command line tool that aggregates search across multiple Chinese music sites (NetEase, QQ Music, Kugou, Baidu, Xiami, Migu) and …
394436maintenance
l429609201/misaka_danmu_server
A self-hosted danmaku (bullet comment) aggregation and management server written in Python, compatible with the dandanplay API specificatio…
811256active
firecrawl/fire-enrich
Fire Enrich is an AI-powered data enrichment web application that transforms a list of email addresses into rich company datasets, includin…
391255active
myreader-io/myGPTReader
myGPTReader is a Slack bot powered by ChatGPT that reads and summarizes webpages, documents (eBooks, PDF, DOCX), and YouTube videos, and su…
604419maintenance
JustinBeckwith/linkinator
Linkinator is a broken link checker that crawls websites, local HTML files, and markdown documentation to find dead links and invalid URLs.…
991254active
MarioVilas/googlesearch
A Python library that performs Google web searches programmatically and returns result URLs, without using an official Google API. It is un…
101252active
minsight-ai-info/AI-Search-Hub
AI Search Hub is an open-source Skill that aggregates native AI search capabilities from platforms like Gemini, Grok, Doubao, and Yuanbao i…
501250active
EmergenceAI/Agent-E
Agent-E is an open-source agent-based system that automates actions on the user's computer, currently focusing on browser automation via na…
611248active
egbertbouman/youtube-comment-downloader
A Python script and library for downloading YouTube video comments without using the official YouTube API. It outputs comments in JSONL, JS…
881247active
AIPexStudio/AIPex
AIPex is an open-source (MIT) AI browser automation assistant delivered as a Chrome/Edge extension that runs inside your existing browser w…
811244active
anvaka/npmgraph.an
A web application that visualizes the dependency graph of npm packages in 2D and 3D, built with Vue 3 and the ngraph library. It fetches pa…
661243active
tomorrow505/auto_feed_js
A Tampermonkey/Greasemonkey userscript that enables one-click cross-posting (re-seeding) of torrents across 100+ private tracker sites. It …
771242active
nonbili/NouTube
NouTube is an Android and desktop application that wraps the mobile YouTube and YouTube Music web apps in a webview, adding ad blocking, ba…
811239active
Decodo/Decodo
Decodo (formerly Smartproxy) is a commercial rotating proxy network and web scraping platform offering 125M+ residential, mobile, ISP, and …
691238active
jeremykendall/php-domain-parser
A PHP library that parses domain names into their component parts (subdomain, registrable domain, second-level domain, public suffix) using…
551238active
seveniruby/AppCrawler
AppCrawler is a Scala-based automated app traversal (crawler) testing tool built on Appium, supporting Android, iOS, mini-programs, and Har…
581236active
raawaa/jav-scrapy
A TypeScript-based Node.js CLI tool that batch-scrapes JAV (adult video) metadata, magnet links, and cover images from source websites. It …
941235active
eatmoreduck/boss-zhipin-scraper
A Python CLI scraper for BOSS Zhipin (zhipin.com) that connects to a locally logged-in Chrome via the Chrome DevTools Protocol to call the …
571231active
nicobailon/pi-web-access
A TypeScript extension for the Pi coding agent that adds web search, URL content extraction, GitHub repo cloning, PDF extraction, and YouTu…
831229active
danny0838/webscrapbook
WebScrapBook is a browser extension that captures web pages faithfully into various archive formats (including MAFF and HTZ) for local or b…
771228active
iszhouhua/social-media-copilot
An open-source browser extension (built with WXT and TypeScript) that scrapes data from Chinese social media platforms including Xiaohongsh…
651228active
saeeddhqan/Maryam
OWASP Maryam is a modular open-source OSINT framework for harvesting data from open sources, search engines, and social networks. It provid…
101228active
ImAiiR/QobuzDownloaderX
QobuzDownloaderX (QBDLX) is a desktop GUI application that downloads music streams directly from the Qobuz streaming platform using its API…
781226active
daijro/browserforge
BrowserForge is a Python library that generates realistic browser headers and fingerprints, mimicking real-world browser, OS, and device di…
561224active
eooce/Auto-login-netlib
A JavaScript automation script that periodically logs into netlib.re accounts to keep free domains alive, run via GitHub Actions on a 60-da…
391224active
firecrawl/web-agent
An open-source TypeScript framework for building autonomous web research agents, layered from a Next.js chat template down to an agent core…
501222active
mika-cn/maoxian-web-clipper
MaoXian Web Clipper is a browser extension for Firefox and Chrome that clips selected content from web pages and saves it to the local mach…
671221active

← prev page 9 / 20 next →