Ross ROSS = Recommend OSS · open-source software intelligence for agents

ssssssss-team/spider-flow

新一代爬虫平台,以图形化方式定义爬虫流程,不写代码即可完成爬虫。 observed · 2026-08-28

github.com/ssssssss-team/spider-flow · homepage · Java · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

23/100

  • Activity 0
  • Release rhythm 8
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2350
  • days_rel: n/a
  • days_push: 1176
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

11352 stars · 2169 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

Spider-Flow is a self-hosted Java-based web crawler platform that lets users define scraping workflows visually as flowcharts without writing code. It supports XPath/JsonPath/CSS selector/regex extraction, JS-rendered pages, proxies, database persistence, and extension via plugins like Selenium, Redis, MongoDB, and OCR.

Use cases

  • scrape websites without writing code
  • build a visual web crawler with a drag-and-drop flowchart
  • extract data from JS-rendered or AJAX pages
  • schedule and monitor crawling tasks with logs
  • save scraped data automatically to a database or files
  • scrape sites behind proxies or captchas with OCR

When to choose

  • you want a no-code/low-code scraping platform with a visual editor
  • your team needs to maintain many crawlers with monitoring and debugging UI
  • you need extensible scraping with plugins for Selenium, proxies, or OCR
  • you want a self-hosted Java crawler platform under an MIT license

When to avoid

  • you need a lightweight code-first scraping library instead of a full platform
  • your project requires active development or recent updates (last release mid-2023)
  • you need distributed large-scale crawling beyond a single self-hosted instance
  • you prefer Python-based scraping ecosystems like Scrapy

Facets

application · maturity maintenance

web-scraping parser workflow-automation plugin-system crawlers web-development developer-tools jvm self-hosted visual-crawler-builder low-code flow-based-programming xpath path css-selectors jsoup selenium proxy-support ocr automation web-server

1 source

Member repositories

RepositoryRoleHealth v2
ssssssss-team/spider-flowmain23

For agents

markdown · JSON · MCP: product_card(name="ssssssss-team/spider-flow")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem