# zorlan/skycaiji

蓝天采集器是一款开源免费的爬虫系统，仅需点选编辑规则即可采集数据，可运行在本地、虚拟主机或云服务器中，几乎能采集所有类型的网页，无缝对接各类CMS建站程序，免登录实时发布数据，全自动无需人工干预！是网页大数据采集软件中完全跨平台的云端爬虫系统

Repository: https://github.com/zorlan/skycaiji
Canonical: https://ross.abutalabs.com/products/skycaiji
Homepage: https://www.skycaiji.com
Language: PHP
License: NOASSERTION
License Family: other
Topics: crawler, crawling, spider, webcrawler, php
Last push: 2026-08-23T06:36:13+00:00

## Health v2 (maintenance only)
Score: 77/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 99, release rhythm 35, longevity 100
- inputs: {"age_days": 3110, "days_push": 10, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, no_license
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2089, forks 608 (observed 2026-08-28T04:06:12.294328+00:00)

## What it is
SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a browser. It supports multi-level/paginated crawling, browser-rendered pages, regex/XPath/JSON extraction, and automatic publishing to CMS platforms, databases, files, or API endpoints.

## Use cases
- scrape website data without coding
- collect articles and publish to a CMS automatically
- extract data from paginated multi-level pages
- gather training data for LLM/AIGC applications
- schedule automated recurring web scraping tasks
- export scraped data to Excel or a database
- scrape JavaScript-rendered pages with a simulated browser
- serve collected data through a REST API

## When to choose
- you want a self-hosted, browser-based visual scraper with no coding
- you need to publish scraped content directly into CMS platforms
- you run PHP hosting or shared virtual hosts and want a cross-platform crawler
- you need scheduled, fully automated collection with proxy rotation

## When to avoid
- you need a programmable scraping library to embed in your own code
- your stack is Python/Node and you prefer code-first frameworks like Scrapy or Playwright
- you need large-scale distributed crawling beyond a single PHP server
- you require a license without restrictions (license is non-standard)

## Facets
- artifact type: application
- maturity: active
- function: web-scraping, etl, parser, workflow-automation, scheduling, plugin-system
- domain: crawlers, web-development, cms
- platform: php, self-hosted, windows, cross-platform
- tags: visual-scraping, no-code-crawler, cms-integration, proxy-rotation, browser-rendering, data-publishing, cloud-deployable, data-engineering, automation, web-server, linux, docker

## Member repositories
- zorlan/skycaiji (main) score 77

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:12.294328+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:55:40.345409+00:00, confidence not recorded.
  - readme: https://github.com/zorlan/skycaiji (fetched 2026-08-28T04:06:12.294328+00:00, sha 8d208bafbce3)
  - homepage: https://www.skycaiji.com (fetched 2026-08-29T10:35:32.119107+00:00, sha 611a63a966e6)
  - site_page: https://www.skycaiji.com/manual/doc/install (fetched 2026-08-29T10:35:32.121750+00:00, sha 68dcbfdb04bb)
- Data as of 2026-08-30T08:39:29.467469+00:00.
