ssssssss-team/spider-flow
新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。 observed · 2026-08-28
Health v2 · maintenance only
23/100
- Activity 0
- Release rhythm 8
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 2350
- days_rel: n/a
- days_push: 1176
- n_releases_24m: 0
Adoption not part of the score
11352 stars · 2169 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
Spider-Flow is a self-hosted Java-based web crawler platform that lets users define scraping workflows visually as flowcharts without writing code. It supports XPath/JsonPath/CSS selector/regex extraction, JS-rendered pages, proxies, database persistence, and extension via plugins like Selenium, Redis, MongoDB, and OCR.
Use cases
- scrape websites without writing code
- build a visual web crawler with a drag-and-drop flowchart
- extract data from JS-rendered or AJAX pages
- schedule and monitor crawling tasks with logs
- save scraped data automatically to a database or files
- scrape sites behind proxies or captchas with OCR
When to choose
- you want a no-code/low-code scraping platform with a visual editor
- your team needs to maintain many crawlers with monitoring and debugging UI
- you need extensible scraping with plugins for Selenium, proxies, or OCR
- you want a self-hosted Java crawler platform under an MIT license
When to avoid
- you need a lightweight code-first scraping library instead of a full platform
- your project requires active development or recent updates (last release mid-2023)
- you need distributed large-scale crawling beyond a single self-hosted instance
- you prefer Python-based scraping ecosystems like Scrapy
Facets
application · maturity maintenance
web-scraping parser workflow-automation plugin-system crawlers web-development developer-tools jvm self-hosted visual-crawler-builder low-code flow-based-programming xpath path css-selectors jsoup selenium proxy-support ocr automation web-server
1 source
- readme: https://github.com/ssssssss-team/spider-flow · fetched 2026-08-28 · d37f324fd5cd
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| ssssssss-team/spider-flow | main | 23 |
For agents
markdown · JSON · MCP: product_card(name="ssssssss-team/spider-flow")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem