Ross ROSS = Recommend OSS · open-source software intelligence for agents

zorlan/skycaiji

蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统 observed · 2026-08-28

github.com/zorlan/skycaiji · homepage · PHP · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

77/100

  • Activity 99
  • Release rhythm 35
  • Longevity 100

Flags: no_releases no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 3110
  • days_rel: n/a
  • days_push: 10
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2089 stars · 608 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a browser. It supports multi-level/paginated crawling, browser-rendered pages, regex/XPath/JSON extraction, and automatic publishing to CMS platforms, databases, files, or API endpoints.

Use cases

  • scrape website data without coding
  • collect articles and publish to a CMS automatically
  • extract data from paginated multi-level pages
  • gather training data for LLM/AIGC applications
  • schedule automated recurring web scraping tasks
  • export scraped data to Excel or a database
  • scrape JavaScript-rendered pages with a simulated browser
  • serve collected data through a REST API

When to choose

  • you want a self-hosted, browser-based visual scraper with no coding
  • you need to publish scraped content directly into CMS platforms
  • you run PHP hosting or shared virtual hosts and want a cross-platform crawler
  • you need scheduled, fully automated collection with proxy rotation

When to avoid

  • you need a programmable scraping library to embed in your own code
  • your stack is Python/Node and you prefer code-first frameworks like Scrapy or Playwright
  • you need large-scale distributed crawling beyond a single PHP server
  • you require a license without restrictions (license is non-standard)

Facets

application · maturity active

web-scraping etl parser workflow-automation scheduling plugin-system crawlers web-development cms php self-hosted windows cross-platform visual-scraping no-code-crawler cms-integration proxy-rotation browser-rendering data-publishing cloud-deployable data-engineering automation web-server linux docker

3 sources

Member repositories

RepositoryRoleHealth v2
zorlan/skycaijimain77

For agents

markdown · JSON · MCP: product_card(name="zorlan/skycaiji")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem