Ross ROSS = Recommend OSS · open-source software intelligence for agents

function: web-scraping

1985 products, primary matches first, then adoption-weighted; health v2 shown.

ProductHealth v2StarsMaturity
crawlab-team/crawlab
Crawlab is a Go-based distributed web crawler management platform with a web UI for managing, scheduling, and monitoring spiders written in…
5212262active
resume/resume.github.com
GitHub Résumé is a hosted web service that automatically generates a résumé page from a user's public GitHub repositories and activity. It …
3262887maintenance
jina-ai/reader
Jina AI Reader converts any URL into LLM-friendly markdown via the r.jina.ai prefix, and searches the web into markdown via s.jina.ai. It r…
6211912active
seanmonstar/reqwest
Reqwest is an ergonomic, batteries-included HTTP client library for Rust built on hyper. It supports async and blocking clients, JSON/multi…
9311800stable
code4craft/webmagic
WebMagic is a scalable web crawler framework for Java covering the full crawl lifecycle: downloading, URL management, content extraction (X…
6011678active
go-shiori/shiori
Shiori is a simple bookmark manager written in Go, intended as a self-hosted clone of Pocket. It works as both a command-line application a…
7411615active
daijro/camoufox
Camoufox is an open-source anti-detect browser built on Firefox for web scraping and AI agents, with browser fingerprint spoofing and Playw…
9011455active
mozilla/readability
Mozilla's standalone Readability library that extracts the main article content, title, and metadata from HTML documents, powering Firefox …
7511410stable
jhy/jsoup
jsoup is a Java library for parsing, manipulating, and cleaning real-world HTML and XML, implementing the WHATWG HTML5 specification to pro…
9511387stable
ihmily/DouyinLiveRecorder
A Python-based livestream recording application that can run unattended in a loop and record multiple streams simultaneously from 40+ platf…
6410786active
blacklanternsecurity/bbot
BBOT is a recursive, multipurpose internet scanner built in Python for automating OSINT reconnaissance, bug bounty hunting, and attack surf…
9910508active
Ranchero-Software/NetNewsWire
NetNewsWire is a free and open-source RSS/Atom/JSON Feed reader for macOS and iOS. It displays articles from blogs and news sites, tracks r…
9910319active
pinchtab/pinchtab
PinchTab is a standalone Go HTTP server that gives AI agents direct control over Chrome via CDP, with stealth injection and multi-instance …
8210144active
shmilylty/OneForAll
OneForAll is a powerful Python-based subdomain collection and enumeration tool for reconnaissance. It gathers subdomains via brute forcing,…
5910029active
thewhiteh4t/seeker
Seeker is a security testing tool that hosts a fake website asking users for browser location permission, capturing GPS coordinates (longit…
729937active
dataabc/weiboSpider
A Python command-line crawler that scrapes posts and profile data from Sina Weibo users, writing results to txt/csv/json files or MySQL/Mon…
629695active
cooderl/wewe-rss
WeWe RSS is a self-hostable service that generates RSS feeds for WeChat official accounts by leveraging WeRead as a data source. It support…
109671active
anvaka/city-roads
A web application that renders every road in any city at once as an interactive visualization, using data from OpenStreetMap. It includes a…
659566active
agefanscom/website
A release page that publishes the current official website URLs and app download links for AGE Animation (AGE动漫), an anime streaming site w…
739540active
ntegrals/openbrowser
Open Browser is a TypeScript framework that lets AI agents autonomously control a web browser to complete tasks like clicking, typing, navi…
669515active
jiji262/douyin-downloader
A Python-based Douyin (TikTok China) downloader that fetches videos, image galleries, collections, and music without watermarks, supporting…
969507active
zubair-trabzada/geo-seo-claude
A Claude Code skill that audits and optimizes websites for AI-powered search engines (ChatGPT, Perplexity, Gemini, Google AI Overviews) alo…
609485active
legado (阅读)
Legado (阅读 3.0) is an open-source Android e-book reader that lets users define custom book sources to read web novels and online content. I…
6147040maintenance
ccfddl/ccf-deadlines
A community-maintained tracker of worldwide academic conference deadlines, offering a website portal, WeChat applet, and multiple extension…
779266active
jiangrui1994/CloudSaver
CloudSaver is a self-hosted web application for searching cloud drive (netdisk) resources across multiple sources and transferring them to …
569238active
FongMi/TV
An open-source Android video streaming app built on CatVod that supports both Android TV (leanback) and mobile UIs, with VOD browsing, live…
759223active
RSS-Bridge/rss-bridge
RSS-Bridge is a PHP web application that generates RSS, Atom, and JSON feeds for websites that don't offer one, using hundreds of site-spec…
759190active
mediago-dev/mediago
MediaGo is a cross-platform video downloader that automatically sniffs m3u8/HLS streams and other video resources from web pages, supportin…
819183active
kepano/defuddle
Defuddle is a TypeScript library, CLI, and hosted service that extracts the main content from web pages, removing clutter like comments, si…
879166active
qiye45/wechatDownload
A desktop tool for batch downloading WeChat Official Account (公众号) articles, saving them as html/mhtml/md/pdf/docx/csv and preserving embed…
909105active
Thysrael/Horizon
Horizon is an AI-powered news aggregation platform that fetches content from sources like Hacker News, GitHub, RSS, Reddit, and Telegram, s…
599040active
HerbertHe/iptv-sources
A service that automatically aggregates and updates IPTV m3u playlist sources from multiple public repositories, with EPG data and optional…
718921active
jo-inc/camofox-browser
A self-hosted anti-detection browser server for AI agents, wrapping the Camoufox Firefox fork that spoofs fingerprints at the C++ level. It…
828894active
everywall/ladder
Ladder is a self-hosted Go web proxy that removes CORS and other headers (like CSP) and modifies HTML/CSS/JS in responses, serving as an al…
888880active
ltaoo/wx_channels_download
A small desktop application that downloads videos from WeChat Channels (视频号) by injecting a download button into the WeChat PC client via a…
898878active
RayWangQvQ/BiliBiliToolPro
BiliTool is a .NET-based automated task tool for Bilibili that performs scheduled tasks like daily check-ins, coin donations, lottery parti…
738816active
eze-is/web-access
An Agent Skill that gives AI coding agents (Claude Code, Cursor, Gemini CLI, etc.) full web access capabilities: three-tier channel dispatc…
658745active
Sitoi/dailycheckin
DailyCheckIn is a Python-based daily check-in script collection for Chinese websites and services (Bilibili, Baidu Tieba, iQiyi, V2EX, AcFu…
748690active
hardikvasa/google-images-download
A Python command-line tool that searches and downloads hundreds of images from Google Images to local storage. It uses Selenium with Chrome…
708684active
XiaoYouChR/Ghost-Downloader-3
Ghost Downloader is a cross-platform, open-source download manager built with Python and PySide6/Qt that handles HTTP, BitTorrent/magnet, F…
958660active
tubearchivist/tubearchivist
Tube Archivist is a self-hosted YouTube media server that downloads videos via yt-dlp, indexes them with metadata in Elasticsearch, and ser…
958395active
mherrmann/helium
Helium is a Python library that provides a high-level, human-friendly API for browser automation on top of Selenium, supporting Chrome and …
928323active
kangvcar/InfoSpider
InfoSpider is an open-source Python toolbox that crawls a user's own personal data from dozens of Chinese and international services (email…
588247active
EstrellaXD/Auto_Bangumi
AutoBangumi is a self-hosted, RSS-based automated anime downloading and organizing tool. It parses RSS feeds from sites like Mikan Project,…
968226active
loks666/get_jobs
An AI-powered job application assistant that automatically submits resumes across Chinese job platforms (Boss Zhipin, 51job, Liepin, Zhaopi…
518166active
epi052/feroxbuster
feroxbuster is a fast, recursive content discovery tool written in Rust that performs forced browsing against web servers. It brute-forces …
728035active
freeok/so-novel
So Novel is a Java-based tool for extracting structured content from web pages and exporting it as EPUB, TXT, or PDF ebooks. It offers CLI,…
947943active
alirezamika/autoscraper
AutoScraper is a Python library that automatically learns scraping rules from a URL or HTML content plus a list of sample data you want to …
667904stable
p1ngul1n0/blackbird
Blackbird is a Python CLI OSINT tool that searches for user accounts by username or email across 600+ social networks and platforms, levera…
467861active
adithya-s-k/omniparse
OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into…
497815active
reconurge/flowsint
Flowsint is an open-source, self-hosted OSINT platform for visual, graph-based investigations, providing entity relationship visualization …
877752active
samuelclay/NewsBlur
NewsBlur is an open-source personal RSS feed reader and social news network with intelligence training, full-text search, story archiving, …
987597active
andeya/pholcus
Pholcus is a distributed, high-concurrency web crawler framework written in pure Go. It supports standalone, server, and client modes with …
907577active
mgdm/htmlq
htmlq is a command-line tool, like jq but for HTML, that extracts content from HTML documents using CSS selectors. Written in Rust, it read…
617576stable
Steel Browser
Steel Browser is an open-source browser API and sandbox that manages browser sessions, proxies, stealth, and lifecycle so developers can bu…
897545active
Suwayomi (Tachidesk)
Suwayomi (formerly Tachidesk) is a free, self-hosted manga reader server that is a rewrite of Tachiyomi for desktop platforms. It runs a se…
977514active
metafizzy/infinite-scroll
Infinite Scroll is a JavaScript plugin that automatically loads and appends the next page of content as the user scrolls, avoiding full pag…
277479stable
mvdctop/Movie_Data_Capture
A Python-based local movie organizer that scrapes and renames movie metadata for media servers. It prepares local movie files with correct …
967432active
cv-cat/Spider_XHS
A Python library that reverse-engineers Xiaohongshu (Little Red Book) signature algorithms and wraps the platform's PC, creator, and Pugong…
887423active
symfony/css-selector
A Symfony PHP component that converts CSS selectors into XPath expressions. It lets developers query HTML or XML documents using familiar C…
997421stable
libcpr/cpr
cpr (C++ Requests) is a modern C++ HTTP client library that wraps libcurl with a simple, Python Requests-inspired API. It supports GET/POST…
837417active
berstend/puppeteer-extra
A modular plugin framework for Puppeteer (and Playwright via playwright-extra) that extends headless browser automation with drop-in plugin…
327398active
TheCraigHewitt/seomachine
SEO Machine is a specialized Claude Code workspace of custom commands, AI agents, and Python analysis modules for researching, writing, and…
607381active
NomaDamas/k-skill
A collection of ~80 LLM agent skills tailored for Korean users, installable into coding agents like Claude Code, Codex, and OpenCode. Skill…
657335active
DIYgod/RSSHub-Radar
RSSHub Radar is a browser extension that helps users quickly discover RSS and RSSHub feeds available on the web pages they visit. It simpli…
707312active
WECENG/ticket-purchase
A Python automation tool for snatching tickets on Damai (大麦), China's major event ticketing platform, using Selenium for the web and Appium…
647114active
go-rod/rod
Rod is a high-level Go library that drives Chrome via the Chrome DevTools Protocol for web automation and scraping. It offers both high-lev…
667077active
Momo707577045/m3u8-downloader
A browser-based tool for extracting and downloading m3u8 (HLS) streaming videos by fetching the m3u8 playlist, downloading .ts segments in …
407055active
autoscrape-labs/pydoll
Pydoll is a Python library for automating Chromium-based browsers directly over the Chrome DevTools Protocol, with no WebDriver binary and …
857048active
BrowserMCP/mcp
Browser MCP is an MCP server plus Chrome extension that lets AI applications like Claude, Cursor, VS Code, and Windsurf control the user's …
287019active
hect0x7/JMComic-Crawler-Python
A Python library providing an API client for the JMComic (18comic) site, supporting both web and mobile endpoints, with album downloading, …
976997stable
InternLM/MindSearch
MindSearch is an open-source, LLM-based multi-agent framework that mimics human cognition to perform deep web searches, similar to Perplexi…
276918active
mishushakov/llm-scraper
A TypeScript library that turns any webpage into structured data using LLMs, built on Playwright and the Vercel AI SDK. It supports multipl…
676915active
webclipper/web-clipper
Web Clipper is an open-source browser extension that saves web content to many note-taking platforms such as Notion, OneNote, Obsidian, Bea…
496831active
ScottSloan/Bili23-Downloader
Bili23-Downloader is a free, open-source, cross-platform desktop application for downloading videos from Bilibili. It supports multithreade…
986812active
bda-research/node-crawler
node-crawler is a TypeScript web crawler/spider library for Node.js that fetches pages and provides server-side DOM parsing with automatic …
916799active
wzdnzd/aggregator
A Python-based platform that crawls free proxy nodes from sources like Telegram, GitHub, and search engines, validates their liveness, and …
766754active
VeNoMouS/cloudscraper
A Python library that wraps Requests to bypass Cloudflare's anti-bot protection pages (IUAM), supporting challenge types v1, v2, v3, and Tu…
346726active
subzeroid/instagrapi
instagrapi is a fast Python wrapper for Instagram's unofficial private (mobile) and public web APIs, supporting login with 2FA, session per…
956713active
adbar/trafilatura
Trafilatura is a Python package and command-line tool for crawling the web and extracting main text, metadata, and comments from raw HTML w…
936709active
davidteather/TikTok-Api
An unofficial Python API wrapper for TikTok.com that retrieves trending content, user information, hashtags, and video data without authent…
926593active
SawyerHood/dev-browser
A CLI tool that lets AI agents and developers control a browser by running sandboxed JavaScript scripts against a full Playwright API. It s…
786560active
apurvsinghgautam/robin
Robin is an AI-powered dark web OSINT investigation tool that uses LLMs to refine queries, filter results from dark web search engines, and…
846438active
brightdata/cli
The official Bright Data CLI (npm package @brightdata/cli) providing terminal access to Bright Data's web scraping, search, and structured …
816406active
lexiforest/curl_cffi
curl_cffi is a Python binding for a curl-impersonate fork via cffi, providing an HTTP client that can impersonate browser TLS/JA3, HTTP/2, …
996390active
s0md3v/Arjun
Arjun is a Python command-line tool that discovers hidden HTTP query parameters for URL endpoints using a large dictionary of over 25,000 p…
266385active
chyroc/WechatSogou
A Python library providing a scraping API for WeChat official accounts based on Sogou WeChat search. It lets you look up account info and r…
546377active
manga-download/hakuneko
HakuNeko is a cross-platform desktop application (built on Electron) for downloading manga, anime, and novels from over 1200 websites via p…
566310active
happycola233/tchMaterial-parser
A cross-platform GUI tool that parses and downloads electronic textbook PDFs from China's National Smart Education Platform for Primary and…
956291active
nickscamara/open-deep-research
An open-source clone of OpenAI's Deep Research, built as a Next.js web application. It uses Firecrawl's search and extract APIs to gather l…
206282active
sparklemotion/nokogiri
Nokogiri is a Ruby library for parsing, querying, modifying, and generating XML and HTML documents, built on native parsers like libxml2, l…
966277stable
Python3WebSpider/ProxyPool
A self-hosted proxy pool service that periodically scrapes free proxy sites, stores and scores them in Redis, tests their availability, and…
636243active
wechatsync/Wechatsync
Wechatsync is an open-source Chrome extension that syncs articles from WeChat Official Accounts (or any webpage) to 29+ content platforms l…
616235active
epiral/bb-browser
bb-browser is a TypeScript CLI and MCP server that lets AI agents control your real Chrome browser using your existing login state, exposin…
706128active
langren1353/GM_script
A collection of Tampermonkey/Greasemonkey userscripts, most notably AC-baidu, which removes redirects and ads from Baidu, Sogou, Google, Bi…
726055active
truelockmc/streambert
Streambert is a cross-platform Electron desktop application for streaming and downloading movies, TV series, and anime. It sources video st…
786022active
MontFerret/ferret
Ferret is a declarative, expression-oriented query language (FQL) with an embeddable Go runtime for querying, transforming, and automating …
836008active
hect0x7/JMComic-APK
A GitHub Actions-based automation that periodically checks for updates to the JM Comic (18comic) Android APK and publishes new versions as …
935992active
InfinityLoop1308/PipePipe
PipePipe is an open-source Android app, forked from NewPipe, for browsing and streaming YouTube, BiliBili, and NicoNico without ads, tracke…
995978active
hanydd/BilibiliSponsorBlock
A browser extension that automatically skips sponsored segments, intros/outros, and like-subscribe reminders in Bilibili videos, ported fro…
935962active

← prev page 2 / 20 next →