function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| crawlab-team/crawlab Crawlab is a Go-based distributed web crawler management platform with a web UI for managing, scheduling, and monitoring spiders written in… | 52 | 12262 | active |
| resume/resume.github.com GitHub Résumé is a hosted web service that automatically generates a résumé page from a user's public GitHub repositories and activity. It … | 32 | 62887 | maintenance |
| jina-ai/reader Jina AI Reader converts any URL into LLM-friendly markdown via the r.jina.ai prefix, and searches the web into markdown via s.jina.ai. It r… | 62 | 11912 | active |
| seanmonstar/reqwest Reqwest is an ergonomic, batteries-included HTTP client library for Rust built on hyper. It supports async and blocking clients, JSON/multi… | 93 | 11800 | stable |
| code4craft/webmagic WebMagic is a scalable web crawler framework for Java covering the full crawl lifecycle: downloading, URL management, content extraction (X… | 60 | 11678 | active |
| go-shiori/shiori Shiori is a simple bookmark manager written in Go, intended as a self-hosted clone of Pocket. It works as both a command-line application a… | 74 | 11615 | active |
| daijro/camoufox Camoufox is an open-source anti-detect browser built on Firefox for web scraping and AI agents, with browser fingerprint spoofing and Playw… | 90 | 11455 | active |
| mozilla/readability Mozilla's standalone Readability library that extracts the main article content, title, and metadata from HTML documents, powering Firefox … | 75 | 11410 | stable |
| jhy/jsoup jsoup is a Java library for parsing, manipulating, and cleaning real-world HTML and XML, implementing the WHATWG HTML5 specification to pro… | 95 | 11387 | stable |
| ihmily/DouyinLiveRecorder A Python-based livestream recording application that can run unattended in a loop and record multiple streams simultaneously from 40+ platf… | 64 | 10786 | active |
| blacklanternsecurity/bbot BBOT is a recursive, multipurpose internet scanner built in Python for automating OSINT reconnaissance, bug bounty hunting, and attack surf… | 99 | 10508 | active |
| Ranchero-Software/NetNewsWire NetNewsWire is a free and open-source RSS/Atom/JSON Feed reader for macOS and iOS. It displays articles from blogs and news sites, tracks r… | 99 | 10319 | active |
| pinchtab/pinchtab PinchTab is a standalone Go HTTP server that gives AI agents direct control over Chrome via CDP, with stealth injection and multi-instance … | 82 | 10144 | active |
| shmilylty/OneForAll OneForAll is a powerful Python-based subdomain collection and enumeration tool for reconnaissance. It gathers subdomains via brute forcing,… | 59 | 10029 | active |
| thewhiteh4t/seeker Seeker is a security testing tool that hosts a fake website asking users for browser location permission, capturing GPS coordinates (longit… | 72 | 9937 | active |
| dataabc/weiboSpider A Python command-line crawler that scrapes posts and profile data from Sina Weibo users, writing results to txt/csv/json files or MySQL/Mon… | 62 | 9695 | active |
| cooderl/wewe-rss WeWe RSS is a self-hostable service that generates RSS feeds for WeChat official accounts by leveraging WeRead as a data source. It support… | 10 | 9671 | active |
| anvaka/city-roads A web application that renders every road in any city at once as an interactive visualization, using data from OpenStreetMap. It includes a… | 65 | 9566 | active |
| agefanscom/website A release page that publishes the current official website URLs and app download links for AGE Animation (AGE动漫), an anime streaming site w… | 73 | 9540 | active |
| ntegrals/openbrowser Open Browser is a TypeScript framework that lets AI agents autonomously control a web browser to complete tasks like clicking, typing, navi… | 66 | 9515 | active |
| jiji262/douyin-downloader A Python-based Douyin (TikTok China) downloader that fetches videos, image galleries, collections, and music without watermarks, supporting… | 96 | 9507 | active |
| zubair-trabzada/geo-seo-claude A Claude Code skill that audits and optimizes websites for AI-powered search engines (ChatGPT, Perplexity, Gemini, Google AI Overviews) alo… | 60 | 9485 | active |
| legado (阅读) Legado (阅读 3.0) is an open-source Android e-book reader that lets users define custom book sources to read web novels and online content. I… | 61 | 47040 | maintenance |
| ccfddl/ccf-deadlines A community-maintained tracker of worldwide academic conference deadlines, offering a website portal, WeChat applet, and multiple extension… | 77 | 9266 | active |
| jiangrui1994/CloudSaver CloudSaver is a self-hosted web application for searching cloud drive (netdisk) resources across multiple sources and transferring them to … | 56 | 9238 | active |
| FongMi/TV An open-source Android video streaming app built on CatVod that supports both Android TV (leanback) and mobile UIs, with VOD browsing, live… | 75 | 9223 | active |
| RSS-Bridge/rss-bridge RSS-Bridge is a PHP web application that generates RSS, Atom, and JSON feeds for websites that don't offer one, using hundreds of site-spec… | 75 | 9190 | active |
| mediago-dev/mediago MediaGo is a cross-platform video downloader that automatically sniffs m3u8/HLS streams and other video resources from web pages, supportin… | 81 | 9183 | active |
| kepano/defuddle Defuddle is a TypeScript library, CLI, and hosted service that extracts the main content from web pages, removing clutter like comments, si… | 87 | 9166 | active |
| qiye45/wechatDownload A desktop tool for batch downloading WeChat Official Account (公众号) articles, saving them as html/mhtml/md/pdf/docx/csv and preserving embed… | 90 | 9105 | active |
| Thysrael/Horizon Horizon is an AI-powered news aggregation platform that fetches content from sources like Hacker News, GitHub, RSS, Reddit, and Telegram, s… | 59 | 9040 | active |
| HerbertHe/iptv-sources A service that automatically aggregates and updates IPTV m3u playlist sources from multiple public repositories, with EPG data and optional… | 71 | 8921 | active |
| jo-inc/camofox-browser A self-hosted anti-detection browser server for AI agents, wrapping the Camoufox Firefox fork that spoofs fingerprints at the C++ level. It… | 82 | 8894 | active |
| everywall/ladder Ladder is a self-hosted Go web proxy that removes CORS and other headers (like CSP) and modifies HTML/CSS/JS in responses, serving as an al… | 88 | 8880 | active |
| ltaoo/wx_channels_download A small desktop application that downloads videos from WeChat Channels (视频号) by injecting a download button into the WeChat PC client via a… | 89 | 8878 | active |
| RayWangQvQ/BiliBiliToolPro BiliTool is a .NET-based automated task tool for Bilibili that performs scheduled tasks like daily check-ins, coin donations, lottery parti… | 73 | 8816 | active |
| eze-is/web-access An Agent Skill that gives AI coding agents (Claude Code, Cursor, Gemini CLI, etc.) full web access capabilities: three-tier channel dispatc… | 65 | 8745 | active |
| Sitoi/dailycheckin DailyCheckIn is a Python-based daily check-in script collection for Chinese websites and services (Bilibili, Baidu Tieba, iQiyi, V2EX, AcFu… | 74 | 8690 | active |
| hardikvasa/google-images-download A Python command-line tool that searches and downloads hundreds of images from Google Images to local storage. It uses Selenium with Chrome… | 70 | 8684 | active |
| XiaoYouChR/Ghost-Downloader-3 Ghost Downloader is a cross-platform, open-source download manager built with Python and PySide6/Qt that handles HTTP, BitTorrent/magnet, F… | 95 | 8660 | active |
| tubearchivist/tubearchivist Tube Archivist is a self-hosted YouTube media server that downloads videos via yt-dlp, indexes them with metadata in Elasticsearch, and ser… | 95 | 8395 | active |
| mherrmann/helium Helium is a Python library that provides a high-level, human-friendly API for browser automation on top of Selenium, supporting Chrome and … | 92 | 8323 | active |
| kangvcar/InfoSpider InfoSpider is an open-source Python toolbox that crawls a user's own personal data from dozens of Chinese and international services (email… | 58 | 8247 | active |
| EstrellaXD/Auto_Bangumi AutoBangumi is a self-hosted, RSS-based automated anime downloading and organizing tool. It parses RSS feeds from sites like Mikan Project,… | 96 | 8226 | active |
| loks666/get_jobs An AI-powered job application assistant that automatically submits resumes across Chinese job platforms (Boss Zhipin, 51job, Liepin, Zhaopi… | 51 | 8166 | active |
| epi052/feroxbuster feroxbuster is a fast, recursive content discovery tool written in Rust that performs forced browsing against web servers. It brute-forces … | 72 | 8035 | active |
| freeok/so-novel So Novel is a Java-based tool for extracting structured content from web pages and exporting it as EPUB, TXT, or PDF ebooks. It offers CLI,… | 94 | 7943 | active |
| alirezamika/autoscraper AutoScraper is a Python library that automatically learns scraping rules from a URL or HTML content plus a list of sample data you want to … | 66 | 7904 | stable |
| p1ngul1n0/blackbird Blackbird is a Python CLI OSINT tool that searches for user accounts by username or email across 600+ social networks and platforms, levera… | 46 | 7861 | active |
| adithya-s-k/omniparse OmniParse is a self-hosted ingestion and parsing platform that converts unstructured data (documents, images, audio, video, web pages) into… | 49 | 7815 | active |
| reconurge/flowsint Flowsint is an open-source, self-hosted OSINT platform for visual, graph-based investigations, providing entity relationship visualization … | 87 | 7752 | active |
| samuelclay/NewsBlur NewsBlur is an open-source personal RSS feed reader and social news network with intelligence training, full-text search, story archiving, … | 98 | 7597 | active |
| andeya/pholcus Pholcus is a distributed, high-concurrency web crawler framework written in pure Go. It supports standalone, server, and client modes with … | 90 | 7577 | active |
| mgdm/htmlq htmlq is a command-line tool, like jq but for HTML, that extracts content from HTML documents using CSS selectors. Written in Rust, it read… | 61 | 7576 | stable |
| Steel Browser Steel Browser is an open-source browser API and sandbox that manages browser sessions, proxies, stealth, and lifecycle so developers can bu… | 89 | 7545 | active |
| Suwayomi (Tachidesk) Suwayomi (formerly Tachidesk) is a free, self-hosted manga reader server that is a rewrite of Tachiyomi for desktop platforms. It runs a se… | 97 | 7514 | active |
| metafizzy/infinite-scroll Infinite Scroll is a JavaScript plugin that automatically loads and appends the next page of content as the user scrolls, avoiding full pag… | 27 | 7479 | stable |
| mvdctop/Movie_Data_Capture A Python-based local movie organizer that scrapes and renames movie metadata for media servers. It prepares local movie files with correct … | 96 | 7432 | active |
| cv-cat/Spider_XHS A Python library that reverse-engineers Xiaohongshu (Little Red Book) signature algorithms and wraps the platform's PC, creator, and Pugong… | 88 | 7423 | active |
| symfony/css-selector A Symfony PHP component that converts CSS selectors into XPath expressions. It lets developers query HTML or XML documents using familiar C… | 99 | 7421 | stable |
| libcpr/cpr cpr (C++ Requests) is a modern C++ HTTP client library that wraps libcurl with a simple, Python Requests-inspired API. It supports GET/POST… | 83 | 7417 | active |
| berstend/puppeteer-extra A modular plugin framework for Puppeteer (and Playwright via playwright-extra) that extends headless browser automation with drop-in plugin… | 32 | 7398 | active |
| TheCraigHewitt/seomachine SEO Machine is a specialized Claude Code workspace of custom commands, AI agents, and Python analysis modules for researching, writing, and… | 60 | 7381 | active |
| NomaDamas/k-skill A collection of ~80 LLM agent skills tailored for Korean users, installable into coding agents like Claude Code, Codex, and OpenCode. Skill… | 65 | 7335 | active |
| DIYgod/RSSHub-Radar RSSHub Radar is a browser extension that helps users quickly discover RSS and RSSHub feeds available on the web pages they visit. It simpli… | 70 | 7312 | active |
| WECENG/ticket-purchase A Python automation tool for snatching tickets on Damai (大麦), China's major event ticketing platform, using Selenium for the web and Appium… | 64 | 7114 | active |
| go-rod/rod Rod is a high-level Go library that drives Chrome via the Chrome DevTools Protocol for web automation and scraping. It offers both high-lev… | 66 | 7077 | active |
| Momo707577045/m3u8-downloader A browser-based tool for extracting and downloading m3u8 (HLS) streaming videos by fetching the m3u8 playlist, downloading .ts segments in … | 40 | 7055 | active |
| autoscrape-labs/pydoll Pydoll is a Python library for automating Chromium-based browsers directly over the Chrome DevTools Protocol, with no WebDriver binary and … | 85 | 7048 | active |
| BrowserMCP/mcp Browser MCP is an MCP server plus Chrome extension that lets AI applications like Claude, Cursor, VS Code, and Windsurf control the user's … | 28 | 7019 | active |
| hect0x7/JMComic-Crawler-Python A Python library providing an API client for the JMComic (18comic) site, supporting both web and mobile endpoints, with album downloading, … | 97 | 6997 | stable |
| InternLM/MindSearch MindSearch is an open-source, LLM-based multi-agent framework that mimics human cognition to perform deep web searches, similar to Perplexi… | 27 | 6918 | active |
| mishushakov/llm-scraper A TypeScript library that turns any webpage into structured data using LLMs, built on Playwright and the Vercel AI SDK. It supports multipl… | 67 | 6915 | active |
| webclipper/web-clipper Web Clipper is an open-source browser extension that saves web content to many note-taking platforms such as Notion, OneNote, Obsidian, Bea… | 49 | 6831 | active |
| ScottSloan/Bili23-Downloader Bili23-Downloader is a free, open-source, cross-platform desktop application for downloading videos from Bilibili. It supports multithreade… | 98 | 6812 | active |
| bda-research/node-crawler node-crawler is a TypeScript web crawler/spider library for Node.js that fetches pages and provides server-side DOM parsing with automatic … | 91 | 6799 | active |
| wzdnzd/aggregator A Python-based platform that crawls free proxy nodes from sources like Telegram, GitHub, and search engines, validates their liveness, and … | 76 | 6754 | active |
| VeNoMouS/cloudscraper A Python library that wraps Requests to bypass Cloudflare's anti-bot protection pages (IUAM), supporting challenge types v1, v2, v3, and Tu… | 34 | 6726 | active |
| subzeroid/instagrapi instagrapi is a fast Python wrapper for Instagram's unofficial private (mobile) and public web APIs, supporting login with 2FA, session per… | 95 | 6713 | active |
| adbar/trafilatura Trafilatura is a Python package and command-line tool for crawling the web and extracting main text, metadata, and comments from raw HTML w… | 93 | 6709 | active |
| davidteather/TikTok-Api An unofficial Python API wrapper for TikTok.com that retrieves trending content, user information, hashtags, and video data without authent… | 92 | 6593 | active |
| SawyerHood/dev-browser A CLI tool that lets AI agents and developers control a browser by running sandboxed JavaScript scripts against a full Playwright API. It s… | 78 | 6560 | active |
| apurvsinghgautam/robin Robin is an AI-powered dark web OSINT investigation tool that uses LLMs to refine queries, filter results from dark web search engines, and… | 84 | 6438 | active |
| brightdata/cli The official Bright Data CLI (npm package @brightdata/cli) providing terminal access to Bright Data's web scraping, search, and structured … | 81 | 6406 | active |
| lexiforest/curl_cffi curl_cffi is a Python binding for a curl-impersonate fork via cffi, providing an HTTP client that can impersonate browser TLS/JA3, HTTP/2, … | 99 | 6390 | active |
| s0md3v/Arjun Arjun is a Python command-line tool that discovers hidden HTTP query parameters for URL endpoints using a large dictionary of over 25,000 p… | 26 | 6385 | active |
| chyroc/WechatSogou A Python library providing a scraping API for WeChat official accounts based on Sogou WeChat search. It lets you look up account info and r… | 54 | 6377 | active |
| manga-download/hakuneko HakuNeko is a cross-platform desktop application (built on Electron) for downloading manga, anime, and novels from over 1200 websites via p… | 56 | 6310 | active |
| happycola233/tchMaterial-parser A cross-platform GUI tool that parses and downloads electronic textbook PDFs from China's National Smart Education Platform for Primary and… | 95 | 6291 | active |
| nickscamara/open-deep-research An open-source clone of OpenAI's Deep Research, built as a Next.js web application. It uses Firecrawl's search and extract APIs to gather l… | 20 | 6282 | active |
| sparklemotion/nokogiri Nokogiri is a Ruby library for parsing, querying, modifying, and generating XML and HTML documents, built on native parsers like libxml2, l… | 96 | 6277 | stable |
| Python3WebSpider/ProxyPool A self-hosted proxy pool service that periodically scrapes free proxy sites, stores and scores them in Redis, tests their availability, and… | 63 | 6243 | active |
| wechatsync/Wechatsync Wechatsync is an open-source Chrome extension that syncs articles from WeChat Official Accounts (or any webpage) to 29+ content platforms l… | 61 | 6235 | active |
| epiral/bb-browser bb-browser is a TypeScript CLI and MCP server that lets AI agents control your real Chrome browser using your existing login state, exposin… | 70 | 6128 | active |
| langren1353/GM_script A collection of Tampermonkey/Greasemonkey userscripts, most notably AC-baidu, which removes redirects and ads from Baidu, Sogou, Google, Bi… | 72 | 6055 | active |
| truelockmc/streambert Streambert is a cross-platform Electron desktop application for streaming and downloading movies, TV series, and anime. It sources video st… | 78 | 6022 | active |
| MontFerret/ferret Ferret is a declarative, expression-oriented query language (FQL) with an embeddable Go runtime for querying, transforming, and automating … | 83 | 6008 | active |
| hect0x7/JMComic-APK A GitHub Actions-based automation that periodically checks for updates to the JM Comic (18comic) Android APK and publishes new versions as … | 93 | 5992 | active |
| InfinityLoop1308/PipePipe PipePipe is an open-source Android app, forked from NewPipe, for browsing and streaming YouTube, BiliBili, and NicoNico without ads, tracke… | 99 | 5978 | active |
| hanydd/BilibiliSponsorBlock A browser extension that automatically skips sponsored segments, intros/outros, and like-subscribe reminders in Bilibili videos, ported fro… | 93 | 5962 | active |