function: web-scraping
1985 products, primary matches first, then adoption-weighted; health v2 shown.
| Product | Health v2 | Stars | Maturity |
|---|---|---|---|
| MikeChongCan/scylla Scylla is an intelligent proxy pool service that automatically crawls, validates, and serves proxy IPs via a JSON API and web UI. It integr… | 34 | 4020 | active |
| yacy/yacy_search_server YaCy is a full search engine application in Java that combines a web crawler, a search index server, and a web front-end. It can run standa… | 76 | 4018 | active |
| tonquer/JMComic-qt A cross-platform desktop client for the JMComic (18comic) comic site, built with Python and Qt (PySide6). It supports browsing, searching, … | 97 | 4010 | active |
| snooppr/snoop Snoop is a Python-based OSINT CLI tool that searches for a given username/nickname across ~5400+ websites, with a focus on the CIS region. … | 75 | 4009 | active |
| mahdibland/V2RayAggregator A Python-based automation that aggregates free proxy nodes (Shadowsocks, SSR, Trojan, Vmess) from public sources, deduplicates them, and sp… | 67 | 4008 | active |
| bdvajstudio/javdb The official mobile app for JavDB, an adult video database website cataloging Japanese and Western adult films with search, tagging, rating… | 67 | 3981 | active |
| hardkoded/puppeteer-sharp PuppeteerSharp is a .NET port of the official Node.js Puppeteer API for controlling headless or headful Chrome and Firefox. It supports nav… | 99 | 3915 | active |
| DO-SAY-GO/tdf TDF (Tree Document Format) is a pair of native desktop apps for capturing live web browsing sessions into portable, offline-replayable .tdf… | 66 | 3908 | active |
| binux/pyspider pyspider is a powerful web crawler (spider) system written in Python with a built-in WebUI for script editing, task monitoring, project man… | 10 | 16774 | maintenance |
| DaKheera47/job-ops A self-hosted job search assistant that aggregates searches across 10+ job boards, AI-scores role fit, tailors CVs per posting, and tracks … | 83 | 3884 | active |
| uncle-novel/uncle-novel Uncle Novel is a cross-platform desktop novel/e-book reader application built with Java and OpenJFX, supporting macOS and Windows. It lets … | 56 | 3884 | active |
| pyload/pyload pyLoad is a free, open-source download manager written in pure Python, designed to be lightweight and extensible. It automates downloads fr… | 75 | 3846 | active |
| zly2006/zhihu-plus-plus Zhihu++ is an open-source third-party Android client for the Chinese Q&A platform Zhihu, built in Kotlin, that removes ads, promotional pos… | 85 | 3845 | active |
| JavScraper/Emby.Plugins.JavScraper A metadata scraper plugin for Emby and Jellyfin media servers that fetches Japanese adult movie information (titles, covers, tags, actress … | 23 | 3796 | active |
| omnivore-app/omnivore Omnivore is a complete, open source read-it-later application for saving and reading articles, with highlighting, notes, search, labels, PD… | 88 | 16226 | maintenance |
| kingname/GeneralNewsExtractor GNE (GeneralNewsExtractor) is a Python library that extracts article title, author, publish time, body text, and images from news page HTML… | 81 | 3788 | active |
| pichillilorenzo/flutter_inappwebview A Flutter plugin for embedding inline WebView widgets, running headless WebViews, and opening in-app browser windows (Chrome Custom Tabs / … | 74 | 3761 | active |
| edoardottt/cariddi Cariddi is a fast command-line web crawler written in Go that takes a list of domains, crawls URLs, and scans for endpoints, secrets, API k… | 84 | 3753 | active |
| Tyrrrz/YoutubeDownloader YoutubeDownloader is a cross-platform desktop GUI application for downloading videos, playlists, and channels from YouTube in formats like … | 99 | 16003 | maintenance |
| foamzou/melody Melody is a self-hosted web application ('your music genie') that helps users search for songs across music and video platforms (NetEase, Q… | 65 | 3744 | active |
| Boris-code/feapder feapder is a powerful, easy-to-use Python web crawling framework offering four spider types (AirSpider, Spider, TaskSpider, BatchSpider) fo… | 74 | 3735 | active |
| Manavarya09/design-extract designlang is an open-source CLI (npx designlang <url>) that uses a Playwright-driven headless browser to reverse-engineer any website's de… | 76 | 3710 | active |
| niumoo/bing-wallpaper A Java-based project that automatically fetches Bing's daily 4K wallpaper and archives it with a browsable web gallery and API. It publishe… | 77 | 3685 | active |
| JoeanAmier/TikTokDownloader DouK-Downloader (formerly TikTokDownloader) is an open-source Python tool for downloading and scraping content from Douyin and TikTok, incl… | 71 | 15584 | maintenance |
| TomBursch/kitchenowl KitchenOwl is a self-hosted grocery list and recipe manager with native mobile, web, and desktop apps built in Flutter and a Flask backend.… | 98 | 3647 | active |
| ibnaleem/gosearch GoSearch is a Go-based CLI OSINT tool that searches a username across 300+ websites to find a person's digital footprint, serving as a Sher… | 84 | 3640 | active |
| EhViewer-NekoInverter/EhViewer An Android client app for browsing E-Hentai/ExHentai galleries, forked from EhViewer with a classic Material Design 2 style. It is maintain… | 84 | 3639 | active |
| waditu/tushare TuShare is a Python library for crawling, cleaning, and storing historical and realtime financial data for China stocks and futures. It pro… | 23 | 15366 | maintenance |
| webtorrent/instant.io Instant.io is a web application for streaming file transfer over WebTorrent (BitTorrent over WebRTC), letting users share and download file… | 77 | 3594 | active |
| arxhr007/Aliens_eye Aliens Eye is an AI-powered OSINT CLI tool that scans 840+ social media and web platforms to find accounts associated with a given username… | 94 | 3587 | active |
| oxylabs/google-ai-mode-scraper A code repository of examples for Oxylabs' Google AI Mode Scraper, a commercial API that sends prompts to Google AI Mode and returns parsed… | 61 | 3581 | active |
| zhongbai2333/Tomato-Novel-Downloader A Rust-based downloader for Tomato (Fanqie) novels that exports books as EPUB and other formats, with TUI, Web UI, and legacy CLI interface… | 75 | 3580 | active |
| zc-zhangchen/any-auto-register A multi-platform automated account registration and management system with a Web UI, plugin-based extensibility, and batch registration sup… | 59 | 3578 | active |
| s0md3v/XSStrike XSStrike is a Python command-line Cross Site Scripting (XSS) detection suite that uses hand-written HTML/JavaScript parsers, context analys… | 31 | 15151 | maintenance |
| codelucas/newspaper newspaper3k is a Python 3 library for discovering, downloading, and parsing news articles from websites. It extracts full text, authors, pu… | 66 | 15144 | maintenance |
| elliotgao2/toapi A Python library that turns any website into a JSON API by declaring fields with CSS/XPath selectors. It fetches and parses pages on demand… | 75 | 3555 | active |
| jstrieb/github-stats A tool that generates GitHub profile and repository statistics visualizations as SVG images using GitHub Actions, including data from priva… | 93 | 3542 | active |
| ytmdl ytmdl is a Python CLI tool (with an optional web app) that downloads songs from YouTube as MP3s and enriches them with metadata like artist… | 23 | 3529 | active |
| pjialin/py12306 A Python-based ticket booking assistant for China's 12306 railway system, supporting multi-account, multi-task ticket purchasing with distr… | 64 | 14905 | maintenance |
| thomasdondorf/puppeteer-cluster A Node.js library that manages a pool of Puppeteer-controlled Chromium instances for running browser tasks in parallel. It handles job queu… | 60 | 3514 | stable |
| Gerapy/Gerapy Gerapy is a distributed crawler management framework built on Scrapy, Scrapyd, Django, and Vue.js. It provides a web dashboard for managing… | 63 | 3512 | active |
| stevenschobert/instafeed.js Instafeed.js is a lightweight JavaScript plugin that displays Instagram photos on a website. It fetches images via the Instagram API and re… | 41 | 3508 | active |
| mxschmitt/playwright-go Playwright for Go is a Go library that automates Chromium, Firefox, and WebKit browsers through a single API, supporting both headless and … | 98 | 3481 | active |
| wushuo894/ani-rss ANI-RSS is a self-hosted Java application that automates anime tracking via RSS feeds: it subscribes to shows, downloads torrents through q… | 86 | 3480 | active |
| MZCretin/RollToolsApi RollToolsApi is a free, long-maintained aggregated REST API service that provides commonly used data endpoints for developers, hosted on an… | 73 | 3478 | active |
| DeepSourceCorp/good-first-issue Good First Issue is a web application that curates beginner-friendly issues from popular open-source projects so new developers can make th… | 64 | 3474 | active |
| growchief/growchief GrowChief is an open-source, self-hostable social media automation and outreach tool that automates actions like sending connection request… | 42 | 3466 | active |
| peasoft/NoMoreWalls A Python project that automatically fetches and merges publicly available proxy nodes from the internet and republishes them as Base64 and … | 77 | 3453 | active |
| gautamkrishnar/blog-post-workflow A GitHub Action that automatically fetches your latest blog posts, StackOverflow activity, or YouTube videos via RSS feeds and displays the… | 97 | 3441 | active |
| oxylabs/free-proxy-list A repository promoting Oxylabs' free tier of US datacenter proxies (HTTP/HTTPS/SOCKS5), offering 5 US IPs, 20 concurrent sessions, and 5GB … | 54 | 3434 | active |
| smallfawn/QLScriptPublic A public collection of JavaScript automation scripts for the Qinglong (青龙) task-scheduling panel, primarily performing daily check-ins, sig… | 76 | 3395 | active |
| opsdisk/pagodo pagodo is a Python CLI tool that automates passive Google dork searches by scraping the Google Hacking Database (GHDB) and running those qu… | 49 | 3387 | active |
| H4ckForJob/dirmap Dirmap is an advanced web directory and file scanning tool written in Python, designed to be more powerful than DirBuster, Dirsearch, cansi… | 44 | 3374 | stable |
| white0dew/XiaohongshuSkills A Python CLI tool and agent Skill that automates Xiaohongshu (RED/RedNote) via Chrome DevTools Protocol, supporting publishing posts with i… | 60 | 3366 | active |
| oxylabs/oxylabs-ai-studio-py A Python SDK for Oxylabs AI Studio, a commercial API offering AI-powered web scraping, crawling, search, site mapping, and browser automati… | 59 | 3354 | active |
| oxylabs/google-news-scraper A free Python-based command-line tool from Oxylabs that scrapes Google News articles by topic, exporting headlines, URLs, and publication d… | 64 | 3329 | active |
| Kovah/LinkAce LinkAce is a self-hosted web application for collecting, organizing, and archiving bookmarks with tags, lists, and advanced search. It auto… | 98 | 3328 | active |
| postaddictme/instagram-php-scraper A PHP library that scrapes Instagram via its web version to fetch account information, photos, videos, stories, and comments, with optional… | 33 | 3327 | active |
| oxylabs/amazon-scraper A Python-based free tool and sample code for Oxylabs' Amazon Scraper API that extracts product, search, offer listing, reviews, Q&A, best s… | 71 | 3326 | active |
| tamnd/kage kage is a Go CLI tool that clones websites into browsable offline folders by rendering each page in headless Chrome, snapshotting the final… | 79 | 3322 | active |
| ocsjs/ocsjs OCS (Online Course Script) is a userscript and companion Electron desktop app that automates online course tasks for Chinese university e-l… | 88 | 3314 | active |
| liuzi6612/nav A lightweight, fully static personal navigation/start-page website built with Angular and ng-zorro-antd, with 800+ curated sites built in a… | 74 | 3314 | active |
| psf/requests-html A Python library that combines HTTP requests with intuitive HTML parsing, offering CSS selector and XPath support plus JavaScript rendering… | 23 | 13814 | maintenance |
| internetarchive/heritrix3 Heritrix is the Internet Archive's open-source, extensible, web-scale, archival-quality web crawler written in Java. It crawls websites and… | 95 | 3305 | active |
| oxylabs/chatgpt-scraper Code examples and documentation for Oxylabs' ChatGPT Scraper, a commercial Web Scraper API endpoint that sends prompts to ChatGPT and retur… | 62 | 3301 | active |
| php-curl-class/php-curl-class PHP Curl Class is a PHP library that wraps PHP's cURL extension in a simple object-oriented API for sending HTTP requests. It supports GET,… | 93 | 3297 | active |
| aapatre/Automatic-Udemy-Course-Enroller-GET-PAID-UDEMY-COURSES-for-FREE A Python script that scrapes coupon sites for free Udemy course coupons and automatically enrolls the user in paid courses for free using S… | 58 | 3291 | active |
| Samueli924/chaoxing A Python command-line tool that automatically completes task points (videos, quizzes) on the Chaoxing Xuexitong / Fanya online learning pla… | 78 | 3285 | active |
| apache/nutch Apache Nutch is a highly extensible and scalable open-source web crawler built on Apache Hadoop data structures. It supports batch crawling… | 77 | 3277 | stable |
| iptv-org/epg A Node.js tool that downloads Electronic Program Guide (EPG) data for thousands of TV channels from hundreds of sources and outputs it as X… | 77 | 3250 | active |
| assetnote/kiterunner Kiterunner is a fast content discovery tool written in Go that bruteforces files, folders, and API routes on web servers. It uses a dataset… | 64 | 3247 | stable |
| itsOwen/CyberScraper-2077 CyberScraper 2077 is an AI-powered web scraping application with a Streamlit GUI that uses OpenAI, Gemini, or local Ollama models to intell… | 67 | 3245 | active |
| pydata/pandas-datareader A Python library that extracts data from remote internet sources such as FRED, Fama/French, World Bank, OECD, and Eurostat directly into pa… | 93 | 3238 | active |
| stickerdaniel/linkedin-mcp-server An open-source MCP server that gives AI assistants like Claude and ChatGPT access to LinkedIn data (profiles, companies, jobs, messages) th… | 82 | 3228 | active |
| jsvine/waybackpack Waybackpack is a Python command-line tool that downloads the entire Wayback Machine archive for a given URL, saving every archived snapshot… | 30 | 3228 | active |
| dembrandt/dembrandt Dembrandt is an open-source Node.js CLI that extracts a website's design system—colors, typography, spacing, borders, shadows, motion, comp… | 83 | 3225 | active |
| oxylabs/ai-crawler-py A Python client library for Oxylabs AI Studio's AI-Crawler, a commercial service that crawls websites from a starting URL, uses natural lan… | 61 | 3218 | active |
| lxf746/any-auto-register A desktop application (Mac/Windows, built on Python/FastAPI with an embedded React UI) that automates registration and lifecycle management… | 76 | 3192 | active |
| Bionus/imgbrd-grabber Grabber is a highly customizable imageboard/booru browser and mass downloader that can fetch thousands of images from multiple booru source… | 88 | 3186 | active |
| imraywang/wewrite WeWrite is a Python-based AI agent skill that automates the full WeChat Official Account content pipeline: topic selection from trending ne… | 79 | 3182 | active |
| nilbuild/githunt GitHunt is a React web application and Chrome extension that lets users explore the most starred GitHub projects by week, with filtering by… | 59 | 3182 | active |
| pingc0y/URLFinder URLFinder is a fast, easy-to-use Go CLI tool that extracts JS files, URLs, and sensitive information from web pages, including hidden unaut… | 80 | 3170 | active |
| s0md3v/Photon Photon is a fast Python-based web crawler designed for OSINT (open-source intelligence) tasks. It extracts URLs, emails, social media accou… | 66 | 13146 | maintenance |
| devanshbatham/ParamSpider ParamSpider is a Python CLI tool that mines URLs for a domain or list of domains from web archives (Wayback Machine), filtering out uninter… | 55 | 3160 | active |
| keon/browser-control A small, fast Rust CLI that drives a real browser over the Chrome DevTools Protocol, designed for coding agents and shell pipelines. It exp… | 91 | 3133 | active |
| omkarcloud/google-maps-scraper A desktop application and API for scraping Google Maps business data, extracting 50+ data points including emails, phone numbers, social pr… | 72 | 3130 | active |
| thp/urlwatch urlwatch is a Python CLI tool that monitors webpages for changes and notifies you via e-mail, terminal, or third-party services like Telegr… | 73 | 3128 | active |
| scrapy/scrapyd Scrapyd is a service daemon for deploying and running Scrapy spiders. It lets you upload Scrapy projects and control spiders via a JSON HTT… | 66 | 3101 | stable |
| Pouzor/homelable Homelable is a self-hosted web application for visualizing homelab infrastructure on interactive network, electrical, and 19-inch rack canv… | 81 | 3100 | active |
| leaperone/MultiPost-Extension MultiPost is an open-source browser extension (with a companion desktop app) that publishes text, images, videos, and other content to 10+ … | 85 | 3097 | active |
| Barabama/FreeNodes A Python crawler that aggregates free proxy nodes (v2ray, Clash, vmess, vless, trojan, ss) from public websites and publishes them as auto-… | 73 | 3096 | active |
| axcore/tartube Tartube is a free, open-source GUI front-end for youtube-dl, yt-dlp and other compatible video downloaders, written in Python 3 / Gtk 3. It… | 83 | 3092 | active |
| sergiotapia/magnetissimo Magnetissimo is a self-hosted Elixir/Phoenix web application that crawls popular torrent sites and indexes magnet links into a local Postgr… | 32 | 3091 | active |
| Virtual-Browser/VirtualBrowser VirtualBrowser is a free, open-source anti-fingerprint browser built on Chromium that lets users create and manage multiple isolated browse… | 93 | 3077 | active |
| sabnzbd/sabnzbd SABnzbd is an open-source binary newsreader written in Python that automates Usenet downloads from NZB files, handling downloading, verific… | 99 | 3075 | active |
| asciimoo/hister Hister is a self-hosted, privacy-focused personal search engine that indexes the full contents of visited web pages and local files. It off… | 83 | 3072 | active |
| mtvpls/MoonTVPlus MoonTVPlus is a self-hosted, enhanced fork of MoonTV that aggregates video from multiple sources into a web-based streaming player built wi… | 60 | 3070 | active |
| lackeyjb/playwright-skill An Agent Skill and Claude Code plugin that lets coding agents write and execute Playwright automation scripts on the fly, from simple page … | 82 | 3068 | active |
| symfony/panther Symfony Panther is a PHP library for browser testing and web scraping that drives real browsers (Chrome, Firefox) via the W3C WebDriver pro… | 73 | 3067 | active |
| antimatter15/splat A WebGL-based real-time viewer for 3D Gaussian Splatting scenes, rendering photorealistic navigable 3D environments from photo-derived spla… | 51 | 3065 | active |