Ross ROSS = Recommend OSS · open-source software intelligence for agents

watercrawl/WaterCrawl

Transform Web Content into LLM-Ready Data observed · 2026-08-28

github.com/watercrawl/WaterCrawl · homepage · TypeScript · NOASSERTION (other) observed · 2026-08-28

Health v2 · maintenance only

82/100

  • Activity 98
  • Release rhythm 84
  • Longevity 44

Flags: no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 9
  • age_days: 626
  • days_rel: 105
  • days_push: 16
  • n_releases_24m: 28

Full methodology

Adoption not part of the score

2010 stars · 252 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

WaterCrawl is a self-hostable web application (Python/Django/Scrapy/Celery) that crawls websites and transforms web content into LLM-ready data like markdown, with sitemap generation, web search, JavaScript rendering, and AI-powered processing. It exposes a REST API with official Python client libraries and a hosted cloud offering with free and paid plans.

Use cases

  • convert website content into markdown for LLM training or RAG
  • crawl a site and extract main content without ads and footers
  • generate a sitemap of all URLs on a website
  • scrape pages with JavaScript rendering and screenshots
  • feed crawled web data into a chatbot or RAG pipeline
  • run a self-hosted scraping API with proxy support

When to choose

  • you need LLM-ready structured data from websites
  • you want a self-hosted alternative to hosted scraping APIs like Firecrawl
  • you need sitemap discovery, search, and crawling in one service
  • you want an API-first crawler with a Python client

When to avoid

  • you only need a lightweight one-off scraper script
  • you require a permissive open-source license (license is not standard OSI)
  • you need heavy client-side browser automation beyond rendering
  • you cannot run Docker or don't want a multi-service stack (Django, Celery, MinIO, database)

Facets

application · maturity active

web-scraping search-engine rag llm-inference api-framework self-hosted crawlers large-language-models web-development developer-tools self-hosted python cli crawler scraper html-to-markdown django scrapy celery sitemap-generation javascript-rendering api-service data-extraction retrieval-augmented-generation docker web-server

7 sources

Member repositories

RepositoryRoleHealth v2
watercrawl/WaterCrawlmain82

For agents

markdown · JSON · MCP: product_card(name="watercrawl/WaterCrawl")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem