function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| EtherDream/jsproxy An online web proxy that uses browser Service Workers to intercept and rewrite requests client-side, so the nginx-based server only forward… | 32 | 9327 | maintenance |
| apify/agent-skills A collection of production-grade agent skills from Apify that give AI coding agents (Claude Code, Cursor, Windsurf, Codex, Gemini CLI) expe… | 59 | 2361 | active |
| tavily-ai/tavily-mcp A Model Context Protocol (MCP) server that exposes Tavily's web search, page extraction, site mapping, and crawling tools to AI assistants … | 65 | 2354 | active |
| vitiko98/qobuz-dl A Python command-line tool for searching, exploring, and downloading lossless and Hi-Res music (FLAC/MP3) from Qobuz. It supports interacti… | 37 | 2351 | active |
| lcomplete/huntly Huntly is a self-hosted, AI-powered personal information hub that captures, archives, and organizes web content, RSS feeds, tweets, and hig… | 86 | 2343 | active |
| geekgeekrun/geekgeekrun GeekGeekRun is a free, open-source desktop application (built on Electron, Vue, Puppeteer, SQLite/TypeORM) that automates job hunting on th… | 79 | 2341 | active |
| FriendsOfPHP/Goutte Goutte is a PHP screen scraping and web crawling library providing a simple API to crawl websites and extract data from HTML/XML responses.… | 10 | 9192 | maintenance |
| miraclx/freyr-js Freyr is a Node.js CLI tool that downloads songs from music streaming services like Spotify, Apple Music, and Deezer. It extracts track met… | 67 | 2330 | active |
| erengy/taiga Taiga is an open-source Windows desktop application that automatically detects anime videos you watch on your computer and syncs your progr… | 67 | 2329 | active |
| gxr404/yuque-dl yuque-dl is a Node.js CLI tool that downloads Yuque (语雀) knowledge bases and documents as local Markdown files. It supports batch downloads… | 70 | 2327 | active |
| EricZhu-42/SteamTradingSiteTracker A Steam skin trading market tracker that continuously monitors sell-to-cash (挂刀) ratios across BUFF, IGXE, C5, UUYP, and ECO platforms for … | 32 | 2324 | active |
| easychen/checkchan-dist Check酱 (CheckChan) is a web page content change monitoring tool consisting of a Chrome/Edge browser extension and a self-hostable cloud com… | 32 | 2321 | active |
| dataabc/weibo-search A Python/Scrapy-based crawler that continuously fetches Weibo keyword and hashtag search results, including full post metadata, images, and… | 71 | 2317 | active |
| chaolucky18/xuexitongScript A Tampermonkey userscript (also runnable via browser console) that automates course video playback on the Xuexitong (Chaoxing) learning pla… | 74 | 2311 | active |
| sjdirect/abot Abot is an open source C# web crawler framework built for speed and flexibility, handling multithreading, HTTP requests, scheduling, and li… | 74 | 2310 | active |
| jaywcjlove/github-rank A GitHub user and repository ranking site that publishes global and China leaderboards of GitHub users by followers and repositories by sta… | 98 | 2308 | active |
| 0xMassi/webclaw webclaw is a Rust-based web extraction toolkit that turns any URL into clean, LLM-ready markdown, JSON, or token-optimized text, including … | 77 | 2305 | active |
| NAStool/nas-tools A self-hosted NAS media library management tool that automates organizing, scraping, and syncing media content for media servers like Plex,… | 10 | 9027 | maintenance |
| Anakin-Inc/anakin AnakinScraper OSS is a self-hosted web scraping API written in Go that turns any website into LLM-ready markdown or structured JSON via a s… | 73 | 2302 | active |
| Xiangyu-CAS/xiaohongshu-ops-skill A skill for the OpenClaw agent that turns it into a Xiaohongshu (RedNote) operations assistant, using browser automation (CDP) to analyze f… | 50 | 2291 | active |
| anaskhan96/soup soup is a small Go library for web scraping with an API modeled after Python's BeautifulSoup. It fetches HTML over HTTP and builds a DOM th… | 96 | 2286 | active |
| zenbu-labs/terminal-browser terminal-browser is a real Chromium-based browser that renders inside your terminal using the kitty graphics protocol and Electron's offscr… | 80 | 2285 | active |
| Future-Scholars/paperlib Paperlib is an open-source, cross-platform desktop application for managing academic papers, built with TypeScript and Electron. It scrapes… | 60 | 2274 | active |
| saifyxpro/HeadlessX HeadlessX is a self-hosted browser automation and web scraping platform powered by Camoufox (a C++-patched Firefox) to bypass anti-bot syst… | 76 | 2268 | active |
| ajayyy/DeArrow DeArrow is an open-source browser extension that crowdsources better titles and thumbnails for YouTube videos, replacing sensationalized cl… | 96 | 2259 | active |
| Johnserf-Seed/TikTokDownload A Python CLI tool for batch-downloading Douyin (Chinese TikTok) content without watermarks, including user profile posts, likes, favorites,… | 23 | 8815 | maintenance |
| ying-ck/fanqienovel-downloader A Python tool that downloads novels from Fanqie Novel (fanqienovel.com) by URL or book ID, with search, batch download, update, and backup … | 89 | 2256 | active |
| MinhasKamal/DownGit DownGit is a web tool that creates direct download links for any public GitHub directory or file, packaging them as a zip archive. It is ho… | 54 | 2256 | active |
| coleam00/mcp-crawl4ai-rag An MCP server that combines Crawl4AI web crawling with RAG capabilities backed by a Supabase vector database, exposing tools for AI agents … | 34 | 2245 | active |
| vasu-devs/JustHireMe JustHireMe is a local-first desktop workbench (Tauri frontend, Python backend) that scrapes job postings, ranks role fit against your profi… | 78 | 2236 | stable |
| goclone-dev/goclone Goclone is a Go CLI utility that downloads entire websites to a local directory, preserving relative link structure so the mirrored site ca… | 62 | 2231 | active |
| dou-jiang/codex-console A Python-based integrated console for automating OpenAI/Codex account registration, login, token retrieval, subscription management, and up… | 72 | 2230 | active |
| wabarc/wayback Wayback is an open-source web archiving tool written in Go that captures and preserves web pages via services like Internet Archive, archiv… | 85 | 2227 | active |
| hhursev/recipe-scrapers A Python library for extracting structured recipe data (title, ingredients, instructions, cooking times, images, nutrients) from cooking we… | 98 | 2216 | active |
| VermiIIi0n/fuckZHS A Python 3 automation script that automatically completes Zhihuishu (智慧树) online course videos, including auto-answering pop-up quiz questi… | 45 | 2216 | active |
| Hubs-Foundation/hubs Hubs is an open-source, browser-based multi-user 3D virtual world and social VR platform built with A-Frame, Three.js, and WebXR/WebRTC. It… | 82 | 2214 | active |
| AnotiaWang/deep-research-web-ui A web UI for the dzhng/deep-research project that performs iterative, AI-driven research by combining search engines, web scraping, and LLM… | 66 | 2205 | active |
| ReaJason/xhs A Python SDK that wraps requests to the Xiaohongshu (Little Red Book) web platform for extracting data. It provides a programmatic client f… | 43 | 2202 | active |
| Imangazaliev/DiDOM DiDOM is a fast and simple PHP library for parsing and manipulating HTML and XML documents. It supports loading from strings, files, or URL… | 52 | 2198 | active |
| mampfes/hacs_waste_collection_schedule A Home Assistant custom integration that fetches waste collection schedules from many service providers, ICS/iCal files, or user-defined da… | 98 | 2195 | active |
| opennaslab/kubespider Kubespider is a self-hosted download orchestration system that turns an idle Linux server into a NAS download center. It uses pluggable sou… | 57 | 2190 | active |
| Owez/yark Yark is a Python CLI tool for archiving YouTube channels, downloading videos and accumulating metadata over time with change reports. It in… | 66 | 2184 | active |
| levigross/grequests GRequests is a Go library that wraps net/http with a convenient, Python Requests-style API. It provides helpers for all HTTP verbs, JSON/XM… | 73 | 2182 | active |
| bigintpro/csdn_downloader A Java (Spring Boot + Dubbo) based web service that downloads CSDN resources such as articles and paid/VIP documents without needing points… | 59 | 2179 | active |
| NobyDa/Script A collection of JavaScript scripts and configuration files for iOS proxy tools such as Surge, Quantumult X, Loon, Stash, and Shadowrocket. … | 76 | 8448 | maintenance |
| ericchiang/pup pup is a command line tool for parsing and filtering HTML using CSS selectors, inspired by jq. It reads HTML from stdin, applies selector-b… | 23 | 8435 | maintenance |
| zhzyker/dismap Dismap is a Go-based asset discovery and identification tool that fingerprints web, TCP, UDP, and TLS services using a rule base of 4500+ w… | 23 | 2163 | active |
| AAndyProgram/SCrawler SCrawler is a Windows GUI application that downloads photos and videos from user profiles across many social media and content sites, inclu… | 93 | 2154 | active |
| philss/floki Floki is an Elixir HTML parser that lets you search document nodes using CSS selectors. It supports multiple parsing backends (mochiweb_htm… | 86 | 2149 | stable |
| AaronL725/grok-register A Python toolkit that automates bulk registration of Grok accounts using real Chromium browser automation, with GUI, CLI, and WebUI interfa… | 58 | 2149 | active |
| bit4woo/domain_hunter_pro Domain Hunter Pro is a Burp Suite plugin (Java jar) for automated domain and subdomain collection, web title fetching, and target managemen… | 63 | 2145 | active |
| Rongronggg9/RSS-to-Telegram-Bot A self-hosted Telegram bot that delivers RSS/Atom feed updates to Telegram chats with rich-text formatting and media support. It is multi-u… | 67 | 2141 | active |
| up209d/ResourcesSaverExt A Chrome browser extension that downloads all resources of a website with one click while preserving the original folder structure. It inte… | 37 | 2141 | active |
| iuroc/bilidown Bilidown is a Bilibili video parsing and downloading desktop application supporting 8K video, Hi-Res audio, Dolby Vision, batch parsing, QR… | 83 | 2135 | active |
| pablouser1/ProxiTok ProxiTok is an open-source alternative frontend for TikTok, inspired by Nitter, written in PHP. It proxies all requests to TikTok server-si… | 34 | 2133 | active |
| mgz0227/legado-Harmony Legado (开源阅读) for HarmonyOS is a free, open-source novel and ebook reader application. It supports custom book sources with user-defined sc… | 92 | 2125 | active |
| benvinegar/counterscale Counterscale is a self-hosted, privacy-friendly web analytics tracker and dashboard built on Cloudflare Workers and Workers Analytics Engin… | 79 | 2122 | active |
| cxfksword/jellyfin-plugin-metashark A Jellyfin metadata plugin that scrapes movie and anime metadata primarily from Douban, with TheMovieDb used to fill in missing episode dat… | 97 | 2119 | active |
| feedjira/feedjira Feedjira is a Ruby library for parsing syndication feeds such as RSS and Atom. It supports extensible and custom parsers, letting users add… | 77 | 2103 | active |
| drunkdream/weread-exporter A Python CLI tool that exports books from WeChat Read (微信读书) into epub, pdf, and mobi formats. It hooks into the web reader's Canvas render… | 64 | 2101 | active |
| yuanzl77/IPTV A Python tool that aggregates IPTV live streaming sources daily, validates them with HTTP checks and FFprobe quality probing, and generates… | 69 | 2100 | active |
| chao325/MaoTai_GUIT A Windows GUI/console application for automated flash-sale purchasing (sniping) on JD, Taobao, and Damai, distributed as a ready-to-run EXE… | 72 | 2097 | active |
| vvoovv/blosm Blosm is a Blender addon that imports real-world geodata—OpenStreetMap buildings and roads, Google 3D city tiles, and terrain—with global c… | 66 | 2097 | active |
| ipfs/public-gateway-checker A web application that displays a list of public IPFS gateways and checks whether each is online, including CORS, IPNS, origin isolation, a… | 98 | 2096 | active |
| zorlan/skycaiji SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a … | 77 | 2089 | active |
| hanc00l/nemo_go Nemo is an automated information-gathering platform for penetration testing that integrates common recon tools (Masscan, Nmap, Subfinder, H… | 88 | 2086 | active |
| prajwalch/TorrentSearch TorrentSearch is an Android app that searches torrents across multiple providers simultaneously, with category filters, detailed results, b… | 86 | 2084 | active |
| AzizKpln/Moriarty-Project Moriarty Project is a web-based phone number investigation tool written in Python that gathers information about a given phone number. It a… | 23 | 2074 | active |
| WebReflection/linkedom LinkeDOM is a triple-linked-list based DOM implementation for DOM-less environments like Node.js and Deno, closely following the DOM standa… | 74 | 2070 | active |
| towfiqi/serpbear SerpBear is an open-source, self-hosted search engine position tracking app for monitoring website keyword rankings in Google. It scrapes S… | 83 | 2064 | active |
| lexbor/lexbor Lexbor is a fast, standards-compliant HTML parser and DOM library written in pure C99, with additional modules for CSS, URL, encoding, and … | 82 | 2055 | active |
| oxylabs/how-to-scrape-google-images A Python-based command-line tool that scrapes Google Images search results, including reverse image search based on a provided image URL. I… | 64 | 2055 | active |
| upbit/pixivpy PixivPy is a Python client library for the Pixiv App-API, supporting authenticated access via refresh tokens. It provides methods for searc… | 38 | 2052 | active |
| oxylabs/how-to-scrape-google-flights A Python-based free scraper tool and tutorial for extracting flight data (prices, times, airlines) from Google Flights pages, either direct… | 58 | 2048 | active |
| probberechts/soccerdata A Python library of scrapers that collect soccer data from popular websites like FBref, ESPN, WhoScored, Sofascore, SoFIFA, Understat, Club… | 93 | 2040 | active |
| wbt5/real-url A Python collection of scripts that extracts real streaming URLs (live stream sources) and danmaku (bullet comments) from 59 Chinese and in… | 32 | 7827 | maintenance |
| rubycdp/ferrum Ferrum is a Ruby library providing a clean, high-level API to control Chrome or Chromium via the Chrome DevTools Protocol (CDP), with no Se… | 90 | 2037 | active |
| pt-plugins/PT-Plugin-Plus PT-Plugin-Plus is a Web Extensions browser plugin for Chrome, Edge, and Firefox that streamlines using private tracker (PT) sites, enabling… | 10 | 7814 | maintenance |
| oxylabs/how-to-scrape-amazon-prices A Python-based example repository and free CLI tool for scraping Amazon product prices, best sellers, search results, and deals from depart… | 64 | 2028 | active |
| Ocyss/boss-helper A browser extension (also available as a userscript) that enhances the Boss Zhipin job platform by removing ads, improving the UI, enabling… | 91 | 2027 | active |
| TrianguloY/URLCheck URLCheck is an open-source Android app that acts as an intermediary when opening URLs, letting users inspect, clean, and modify links befor… | 87 | 2027 | active |
| elliotgao2/gain Gain is an asynchronous web crawling framework for Python built on asyncio, aiohttp, and lxml/pyquery. Users declare items and parsers decl… | 75 | 2019 | active |
| afar1/fieldtheory-cli A TypeScript CLI that syncs X/Twitter bookmarks to local markdown files, provides BM25 full-text search and LLM-based classification, and e… | 58 | 2015 | active |
| john-kurkowski/tldextract A Python library that accurately splits URLs into subdomain, domain, and public suffix components using the Public Suffix List. It also shi… | 86 | 2014 | stable |
| jonhoo/fantoccini Fantoccini is a Rust library providing a high-level async API for programmatically controlling browsers via the WebDriver protocol. It supp… | 75 | 2014 | active |
| dmzz-yyhyy/LightNovelReader LightNovelReader is an open-source Android light novel reader app built with Kotlin and Jetpack Compose, featuring a lightweight footprint … | 89 | 2012 | active |
| jimmc414/onefilellm OneFileLLM is a Python command-line tool and library that aggregates content from sources like GitHub repos, pull requests, arXiv/Sci-Hub p… | 69 | 2011 | active |
| watercrawl/WaterCrawl WaterCrawl is a self-hostable web application (Python/Django/Scrapy/Celery) that crawls websites and transforms web content into LLM-ready … | 82 | 2010 | active |
| AnswerOverflow/AnswerOverflow Answer Overflow is an open-source application that indexes Discord server threads into searchable, SEO-friendly web pages so community know… | 65 | 2007 | active |
| supermemoryai/markdowner Markdowner is a fast web service that converts any website into LLM-ready markdown, with optional LLM filtering, detailed responses, and au… | 24 | 1999 | active |
| nottelabs/notte Notte is a full-stack framework and cloud platform for building, deploying, and scaling AI web agents and browser automations. It combines … | 84 | 1997 | active |
| ChinaGodMan/UserScripts A collection of Tampermonkey/Greasyfork userscripts modified from the internet, written in JavaScript. The scripts add browser enhancements… | 68 | 1997 | active |
| zhegexiaohuozi/SeimiCrawler SeimiCrawler is an agile, standalone, distributed Java crawler framework inspired by Python's Scrapy, with deep Spring Boot integration and… | 85 | 1990 | active |
| karpathy/jobs A research tool that scrapes the Bureau of Labor Statistics Occupational Outlook Handbook (342 occupations) and renders an interactive tree… | 47 | 1989 | active |
| A9T9/RPA Ui.Vision RPA is an open-source robotic process automation tool delivered as a browser extension for Chrome, Edge, and Firefox, compatible … | 96 | 1985 | active |
| danmactough/node-feedparser A Node.js library for parsing RSS, Atom, and RDF syndication feeds as a streaming interface. It resolves relative URLs and correctly handle… | 70 | 1976 | stable |
| ViennaRSS/vienna-rss Vienna is a free, open-source RSS/Atom/JSON feed newsreader for macOS with a native Apple Mail-like interface. It can fetch feeds directly … | 97 | 1972 | active |
| walkingddd/TgtoDrive TgtoDrive is a self-hosted, Docker-deployed media automation platform that chains resource discovery (Telegram channel monitoring, search s… | 75 | 1964 | active |
| Flexget/Flexget FlexGet is a multipurpose Python automation tool for content like torrents, NZBs, podcasts, comics, series, and movies. It pulls from sourc… | 95 | 1963 | active |
| Tabula Tabula is a local web application for extracting data tables from text-based PDF files into CSV, Excel, or JSON. It is powered by the tabul… | 28 | 7472 | maintenance |