function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| lxml/lxml lxml is a fast, feature-rich Python library for processing XML and HTML, built on the C libraries libxml2 and libxslt. It provides an Eleme… | 99 | 3050 | stable |
| Qianlitp/crawlergo crawlergo is a Go-based browser crawler that uses headless Chrome to discover URLs for web vulnerability scanners. It renders pages, fills … | 27 | 3034 | active |
| web-scrobbler/web-scrobbler Web Scrobbler is a browser extension that scrobbles music playback history from streaming sites to services like Last.fm, Libre.fm, ListenB… | 90 | 3023 | active |
| symfony/browser-kit Symfony BrowserKit is a PHP component that simulates the behavior of a web browser, enabling programmatic requests, link clicks, and form s… | 99 | 3012 | stable |
| TheBlewish/Automated-AI-Web-Researcher-Ollama A Python application that turns a locally-run Ollama LLM into an automated web research assistant. Given a single query, it generates focus… | 21 | 3012 | active |
| vitalets/github-trending-repos A GitHub-based service that tracks trending repositories per programming language and delivers updates via native GitHub notifications. A s… | 56 | 2995 | active |
| adryfish/fingerprint-chromium A fingerprint browser built on Ungoogled Chromium that lets users spoof or control browser fingerprint characteristics to avoid detection. … | 76 | 2991 | active |
| dw-dengwei/daily-arXiv-ai-enhanced A tool that automatically crawls daily arXiv papers and summarizes them using LLMs like DeepSeek, publishing the results as a browsable sit… | 64 | 2991 | active |
| spacecowboy/Feeder Feeder is an open-source RSS/Atom/JSON Feed reader app for Android, built with Kotlin and Jetpack Compose. It runs entirely locally with no… | 98 | 2988 | active |
| 5ime/video_spider A PHP-based web service that parses short-video links from platforms like Douyin, Kuaishou, Weibo, and Pipixia to return watermark-free vid… | 57 | 2975 | active |
| MALSync/MALSync MAL-Sync is a browser extension and userscript that integrates MyAnimeList, AniList, Kitsu, Simkl, Shikimori, and MangaBaka into hundreds o… | 89 | 2954 | active |
| thewhiteh4t/FinalRecon FinalRecon is an all-in-one automatic web reconnaissance tool written in Python that provides a fast overview of a web target. It bundles h… | 70 | 2953 | active |
| nashsu/AutoCLI AutoCLI is a blazing-fast, memory-safe command-line tool written in Rust that fetches information from 55+ websites (Twitter/X, Reddit, You… | 65 | 2950 | active |
| rust-headless-chrome/rust-headless-chrome A Rust library providing a high-level API to control headless Chrome or Chromium via the DevTools Protocol, serving as the Rust equivalent … | 81 | 2947 | active |
| striver-ing/wechat-spider An open-source WeChat crawler that scrapes articles, reading counts, likes, and comments from WeChat official accounts using a man-in-the-m… | 70 | 2942 | active |
| oxylabs/perplexity-scraper A repository of code examples and documentation for Oxylabs' Perplexity Scraper API, which sends prompts to Perplexity and returns AI-gener… | 62 | 2922 | active |
| RayeRen/acad-homepage.github.io AcadHomepage is a Jekyll-based template for building a modern, responsive academic personal homepage hosted on GitHub Pages. It automatical… | 76 | 2916 | active |
| deedy5/ddgs DDGS (Dux Distributed Global Search) is a Python metasearch library that aggregates text, image, video, news, and book results from multipl… | 95 | 2915 | active |
| Tyrrrz/DiscordChatExporter DiscordChatExporter is an application that exports Discord message history from direct messages, group messages, and server channels to fil… | 95 | 11842 | maintenance |
| ChanceYu/front-end-rss An automated RSS aggregator that collects the latest front-end technology articles from popular newsletters and blogs, categorizes them, an… | 77 | 2894 | active |
| public-clis/twitter-cli A terminal-first CLI for Twitter/X that lets users read timelines, bookmarks, search results, and user profiles without official API keys, … | 51 | 2877 | active |
| cdhigh/KindleEar KindleEar is a self-hostable Python web application that aggregates RSS/ATOM/JSON feeds and web content (including Calibre recipes) into ep… | 73 | 2865 | active |
| zzzprojects/html-agility-pack Html Agility Pack (HAP) is a free, open-source HTML parser written in C# that builds a read/write DOM and supports XPath and XSLT queries. … | 95 | 2846 | active |
| Alfredredbird/tookie-osint Tookie-OSINT is an open-source Python OSINT tool that finds social media accounts and gathers information based on user inputs. It offers a… | 82 | 2841 | active |
| goniszewski/grimoire Grimoire is a local-first, self-hosted bookmark manager for saving, extracting, and searching web content, with optional AI-powered summari… | 96 | 2838 | active |
| spatie/crawler A PHP library by Spatie for crawling links on websites, built on Guzzle promises for concurrent requests. It can execute JavaScript via Chr… | 97 | 2829 | active |
| youniaogu/MangaReader A cross-platform manga reading app built with React Native for Android and iOS, with tablet support. It uses a plugin-based design to aggre… | 51 | 2817 | active |
| cv-cat/DouYin_Spider A Python-based Douyin (Chinese TikTok) reverse-engineering toolkit that exposes the platform's full API surface for data collection, live-s… | 62 | 2812 | active |
| chrieke/prettymapp Prettymapp is a Python package and Streamlit webapp that generates stylized, artistic maps from OpenStreetMap data. It is a speed-focused r… | 82 | 2809 | active |
| digininja/CeWL CeWL is a Ruby command-line tool that spiders a target website to a specified depth and collects unique words into a custom wordlist for us… | 63 | 2801 | stable |
| ssssssss-team/spider-flow Spider-Flow is a self-hosted Java-based web crawler platform that lets users define scraping workflows visually as flowcharts without writi… | 23 | 11352 | maintenance |
| meeb/tubesync TubeSync is a self-hosted PVR (personal video recorder) for YouTube that syncs channels and playlists to local directories, wrapping yt-dlp… | 98 | 2786 | active |
| Geniusay/ChopperBot ChopperBot is a fully automated AI bot that monitors popular live streams across platforms like Douyu, Huya, Bilibili, Douyin, and Twitch, … | 45 | 2779 | active |
| besuper/TwitchNoSub A browser extension for Chrome, Chromium-based browsers, and Firefox that lets users watch subscriber-only VODs (past broadcasts) on Twitch… | 80 | 2776 | active |
| geziyor/geziyor Geziyor is a fast web crawling and scraping framework for Go, supporting JavaScript rendering via Chrome, caching, proxy management, and au… | 73 | 2775 | active |
| rusq/slackdump Slackdump is a Go CLI tool that archives private and public Slack messages, threads, files, users, channels, and emojis locally without req… | 98 | 2764 | active |
| langchain-ai/social-media-agent A TypeScript AI agent built on LangGraph that takes a URL, scrapes its content, and generates Twitter and LinkedIn posts using LLMs. It inc… | 66 | 2756 | active |
| cobrateam/splinter Splinter is a Python library providing a simple, consistent API for browser automation and web application acceptance testing. It abstracts… | 39 | 2753 | active |
| GodsScion/Auto_job_applier_linkedIn A Python/Selenium desktop application that automates LinkedIn Easy Apply job applications, finding relevant jobs, filling application quest… | 72 | 2736 | active |
| xnl-h4ck3r/waymore waymore is a Python CLI tool that retrieves URLs from multiple web archive and intelligence sources (Wayback Machine, Common Crawl, Alien V… | 90 | 2732 | active |
| microlinkhq/metascraper Metascraper is a Node.js library that extracts unified metadata from any URL by combining Open Graph, JSON-LD, Microdata, RDFa, Twitter Car… | 95 | 2731 | active |
| vladkens/twscrape twscrape is an async Python library and CLI for scraping X/Twitter via its Search and GraphQL endpoints using a pool of your own accounts. … | 96 | 2708 | active |
| Dineshkarthik/telegram_media_downloader A Python tool that downloads media files (audio, documents, photos, videos, voice notes) from Telegram chats, channels, and conversations, … | 81 | 2707 | active |
| Nandaka/PixivUtil2 A Python command-line tool for bulk downloading images from Pixiv and Pixiv FANBOX, with support for downloading by member, tag, bookmark, … | 71 | 2698 | active |
| jae-jae/QueryList QueryList is a progressive PHP web scraping framework built on phpQuery that provides jQuery-like CSS3 DOM selectors and manipulation APIs … | 74 | 2690 | active |
| Open-Web-Analytics/Open-Web-Analytics Open Web Analytics (OWA) is an open source, self-hosted web analytics server and JavaScript tracking client, serving as an alternative to G… | 99 | 2685 | active |
| CharlesPikachu/videodl A lightweight video downloader written in pure Python that parses and downloads videos from dozens of streaming platforms (Douyin, Bilibili… | 83 | 2676 | active |
| chrome-php/chrome A PHP library for controlling headless Chrome/Chromium browsers via the DevTools protocol, supporting both synchronous and asynchronous usa… | 89 | 2675 | active |
| spider-rs/spider Spider is a concurrency-first web crawler and scraper written in Rust that streams pages as they arrive, renders JavaScript only when neede… | 87 | 2672 | active |
| ShareDropio/sharedrop ShareDrop is a web application inspired by Apple AirDrop that transfers files directly between devices using WebRTC peer-to-peer connection… | 35 | 10749 | maintenance |
| prinsss/twitter-web-exporter A UserScript (for Tampermonkey/Violentmonkey) that exports tweets, bookmarks, lists, followers, search results, and media from the Twitter/… | 86 | 2666 | active |
| LittleSurvival/copymanga-copy20 A collection of Chinese manga source extensions for Tachiyomi/Mihon-compatible comic reader apps, supporting Copymanga, vomic, Baozimanhua,… | 86 | 2636 | active |
| TermuxHackz/X-osint X-osint is an open-source Python-based OSINT framework for gathering information about phone numbers, email addresses, IP addresses, VINs, … | 72 | 2627 | active |
| Johnserf-Seed/f2 F2 is an asynchronous Python library and CLI tool for downloading videos and fetching API data from multiple platforms including Douyin, Ti… | 56 | 2620 | active |
| spatie/laravel-sitemap A Laravel package by Spatie that generates XML sitemaps, either by crawling an entire site automatically or by adding URLs manually (includ… | 95 | 2617 | stable |
| MrsEWE44/musicDownload A Python GUI application for searching and downloading music from major Chinese music platforms (Kugou, Kuwo, QQ Music, NetEase Cloud Music… | 69 | 2611 | active |
| brightdata/brightdata-mcp A Model Context Protocol (MCP) server by Bright Data that gives AI agents and LLMs real-time access to public web data through 69 tools cov… | 84 | 2610 | active |
| raviqqe/muffet Muffet is a fast website link checker written in Go that recursively scrapes and inspects all pages of a website for broken links. It suppo… | 91 | 2609 | active |
| zhizhuodemao/js-reverse-mcp An MCP server that gives AI coding agents (Claude, Cursor, Copilot) tools for JavaScript reverse engineering in a headed Chrome browser, in… | 83 | 2608 | active |
| Serene-Arc/bulk-downloader-for-reddit A Python command-line tool (bdfr) that bulk-downloads and archives Reddit submissions and their media from subreddits, multireddits, users,… | 57 | 2606 | active |
| botswin/BotBrowser BotBrowser is a privacy-focused browser core (Chromium-based) that unifies and controls browser fingerprint signals across platforms, integ… | 85 | 2589 | active |
| liu673cn/bug An open-source TVBox shell application (empty shell) for Android TV and mobile devices that plays video streams from user-supplied configur… | 32 | 10354 | maintenance |
| apify/fingerprint-suite A modular TypeScript toolkit by Apify for generating realistic browser fingerprints and HTTP headers and injecting them into Playwright or … | 98 | 2581 | active |
| sarperavci/CloudflareBypassForScraping A Python library that bypasses Cloudflare's anti-bot verification for web scraping, supporting cookie generation and request mirroring for … | 72 | 2576 | active |
| miantiao-me/hacker-podcast An AI-powered Chinese podcast application that automatically scrapes daily Hacker News top stories, generates Chinese summaries and scripts… | 62 | 2576 | active |
| lncrawl/lightnovel-crawler Lightnovel Crawler is a Python tool that downloads web novels from 300+ supported sources and converts them into e-books such as EPUB, MOBI… | 97 | 2575 | active |
| amtoaer/bili-sync bili-sync is a Bilibili video synchronization tool written in Rust and Tokio, designed for NAS users. It automatically downloads favorites,… | 88 | 2570 | active |
| pt-plugins/PT-depiler PT-depiler is a Manifest v3 browser extension that improves efficiency when using private tracker (PT) sites, succeeding PT-Plugin-Plus. It… | 99 | 2566 | active |
| simonw/shot-scraper shot-scraper is a Python CLI utility built on Playwright for taking automated screenshots of websites, recording video demos, and scraping … | 89 | 2553 | active |
| zubair-trabzada/ai-marketing-claude A suite of 15 marketing skills for Claude Code that runs parallel subagents to audit websites, generate copy, email sequences, ad creatives… | 46 | 2551 | active |
| owntone/owntone-server OwnTone is an open-source audio media server written in C that streams local files, Spotify, and internet radio to AirPlay 1/2 (multiroom),… | 93 | 2544 | active |
| get-iplayer/get_iplayer A Perl-based command-line tool that indexes and downloads TV and radio programmes from BBC iPlayer and BBC Sounds. It supports regex search… | 31 | 2538 | active |
| jackwener/xiaohongshu-cli A Python CLI for Xiaohongshu (小红书/Red) that lets users search, read, and interact with notes, comments, feeds, and profiles via a reverse-e… | 48 | 2534 | active |
| nobiyou/wx_channel A Windows desktop tool that injects download buttons into WeChat Channels (视频号) pages, enabling one-click and batch downloading of videos i… | 86 | 2531 | active |
| BetterBahn/betterbahn BetterBahn is an open-source Next.js web app for finding train journeys in Germany, with a focus on split-ticketing to help users save mone… | 62 | 2529 | active |
| zotify-dev/zotify Zotify is a command-line tool for downloading music tracks, albums, playlists, and podcasts directly from streaming sources at up to 320kbp… | 32 | 2523 | active |
| guyueyingmu/avbook A self-hosted PHP/Laravel web application that manages a Japanese adult video (JAV) library, backed by crawlers for sites like avmoo, javbu… | 23 | 10036 | maintenance |
| kpcyrd/sn0int sn0int is a semi-automatic OSINT framework and package manager written in Rust that enumerates attack surface by processing public informat… | 60 | 2515 | active |
| m4ll0k/SecretFinder SecretFinder is a Python CLI script based on LinkFinder that discovers sensitive data like API keys, access tokens, and JWTs in JavaScript … | 32 | 2500 | active |
| krau/SaveAny-Bot A self-hosted Telegram bot written in Go that saves Telegram files (documents, videos, photos, stickers, Telegraph pages) to various storag… | 85 | 2491 | active |
| tid-kijyun/Kanna Kanna is a Swift XML/HTML parser library inspired by Ruby's Nokogiri, built on libxml2. It supports XPath 1.0 and CSS3 selector queries acr… | 72 | 2486 | active |
| Luoyacheng/legado-E 阅读Sigma (Reading Sigma) is an open-source Android e-book reader forked from Legado, offering customizable book sources, RSS subscriptions, … | 74 | 2485 | active |
| fhamborg/news-please news-please is an open-source Python news crawler and information extractor that pulls structured article data (headline, lead, main text, … | 67 | 2482 | active |
| iliane5/meridian Meridian is an open-source AI-powered news intelligence system that scrapes hundreds of RSS sources, clusters and analyzes articles with LL… | 31 | 2442 | active |
| coursera-dl/coursera-dl A Python command-line script for batch downloading Coursera.org lecture videos and resources with proper naming. It supports resuming downl… | 32 | 9645 | maintenance |
| badlogic/pi-skills A collection of skills (SKILL.md-based tool packages) for the pi coding agent, also compatible with Claude Code, Codex CLI, Amp, and Droid.… | 55 | 2434 | active |
| dimdenGD/OldTweetDeck A browser extension that restores the classic (old) TweetDeck interface on X/Twitter. It works on Chromium-based browsers and Firefox by in… | 75 | 2429 | active |
| rust-scraper/scraper A Rust library for parsing HTML documents and querying them with CSS selectors, built on Servo's html5ever and selectors crates for browser… | 89 | 2417 | active |
| kurtmckee/feedparser feedparser is a Python library that parses RSS, Atom, RDF, CDF, and JSON feeds into a uniform data structure. It is the de facto standard u… | 88 | 2414 | stable |
| lmnr-ai/index Index is an open-source browser agent that autonomously performs complex tasks on the web by turning any website into an accessible API. It… | 10 | 2411 | active |
| scrapinghub/portia Portia is a visual web scraping tool from Scrapinghub built on Scrapy that lets users annotate web pages in a browser to define data extrac… | 10 | 9504 | maintenance |
| ccloli/E-Hentai-Downloader A userscript that downloads E-Hentai/ExHentai galleries as zip files directly from the browser. It runs via userscript managers like Tamper… | 60 | 2404 | active |
| jarvis2f/telegram-files A self-hosted web application for downloading files from Telegram channels and groups continuously and unattended. It supports multiple Tel… | 88 | 2389 | active |
| OpenBullet OpenBullet 2 is a cross-platform automation suite built on .NET for performing HTTP requests against target web applications and processing… | 85 | 2389 | active |
| Bistutu/GoMusic GoMusic is a web application that migrates playlists from Chinese music services (NetEase Cloud, Qishui, QQ Music) to Apple Music, YouTube … | 59 | 2386 | active |
| gawel/pyquery pyquery is a Python library that provides a jQuery-like API for querying and manipulating XML and HTML documents, built on top of lxml for … | 75 | 2377 | stable |
| hisxo/gitGraber gitGraber is a Python3 command-line tool that monitors GitHub search results in real time to find leaked sensitive data such as API keys an… | 66 | 2376 | active |
| kevinho/clawfeed ClawFeed is an AI-powered news digest application that curates sources like Twitter, RSS, HackerNews, Reddit, and GitHub Trending into stru… | 64 | 2376 | active |
| StractOrg/stract Stract is an open-source web search engine written in Rust with its own independent crawler and index, built on the Tantivy inverted index … | 10 | 2373 | active |
| lijiejie/BBScan BBScan is a fast, lightweight, high-concurrency web vulnerability scanner written in Python. It helps penetration testers quickly identify … | 23 | 2371 | active |