zorlan/skycaiji
蓝天采集器是一款开源免费的爬虫系统,仅需点选编辑规则即可采集数据,可运行在本地、虚拟主机或云服务器中,几乎能采集所有类型的网页,无缝对接各类CMS建站程序,免登录实时发布数据,全自动无需人工干预!是网页大数据采集软件中完全跨平台的云端爬虫系统 observed · 2026-08-28
Health v2 · maintenance only
77/100
- Activity 99
- Release rhythm 35
- Longevity 100
Flags: no_releases no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 3110
- days_rel: n/a
- days_push: 10
- n_releases_24m: 0
Adoption not part of the score
2089 stars · 608 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
SkyCaiji (蓝天采集器) is an open-source, PHP+MySQL based visual web scraping system where users define collection rules by point-and-click in a browser. It supports multi-level/paginated crawling, browser-rendered pages, regex/XPath/JSON extraction, and automatic publishing to CMS platforms, databases, files, or API endpoints.
Use cases
- scrape website data without coding
- collect articles and publish to a CMS automatically
- extract data from paginated multi-level pages
- gather training data for LLM/AIGC applications
- schedule automated recurring web scraping tasks
- export scraped data to Excel or a database
- scrape JavaScript-rendered pages with a simulated browser
- serve collected data through a REST API
When to choose
- you want a self-hosted, browser-based visual scraper with no coding
- you need to publish scraped content directly into CMS platforms
- you run PHP hosting or shared virtual hosts and want a cross-platform crawler
- you need scheduled, fully automated collection with proxy rotation
When to avoid
- you need a programmable scraping library to embed in your own code
- your stack is Python/Node and you prefer code-first frameworks like Scrapy or Playwright
- you need large-scale distributed crawling beyond a single PHP server
- you require a license without restrictions (license is non-standard)
Facets
application · maturity active
web-scraping etl parser workflow-automation scheduling plugin-system crawlers web-development cms php self-hosted windows cross-platform visual-scraping no-code-crawler cms-integration proxy-rotation browser-rendering data-publishing cloud-deployable data-engineering automation web-server linux docker
3 sources
- readme: https://github.com/zorlan/skycaiji · fetched 2026-08-28 · 8d208bafbce3
- homepage: https://www.skycaiji.com · fetched 2026-08-29 · 611a63a966e6
- site_page: https://www.skycaiji.com/manual/doc/install · fetched 2026-08-29 · 68dcbfdb04bb
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| zorlan/skycaiji | main | 77 |
For agents
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem