watercrawl/WaterCrawl
Transform Web Content into LLM-Ready Data observed · 2026-08-28
Health v2 · maintenance only
82/100
- Activity 98
- Release rhythm 84
- Longevity 44
Flags: no_license
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 9
- age_days: 626
- days_rel: 105
- days_push: 16
- n_releases_24m: 28
Adoption not part of the score
2010 stars · 252 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
WaterCrawl is a self-hostable web application (Python/Django/Scrapy/Celery) that crawls websites and transforms web content into LLM-ready data like markdown, with sitemap generation, web search, JavaScript rendering, and AI-powered processing. It exposes a REST API with official Python client libraries and a hosted cloud offering with free and paid plans.
Use cases
- convert website content into markdown for LLM training or RAG
- crawl a site and extract main content without ads and footers
- generate a sitemap of all URLs on a website
- scrape pages with JavaScript rendering and screenshots
- feed crawled web data into a chatbot or RAG pipeline
- run a self-hosted scraping API with proxy support
When to choose
- you need LLM-ready structured data from websites
- you want a self-hosted alternative to hosted scraping APIs like Firecrawl
- you need sitemap discovery, search, and crawling in one service
- you want an API-first crawler with a Python client
When to avoid
- you only need a lightweight one-off scraper script
- you require a permissive open-source license (license is not standard OSI)
- you need heavy client-side browser automation beyond rendering
- you cannot run Docker or don't want a multi-service stack (Django, Celery, MinIO, database)
Facets
application · maturity active
web-scraping search-engine rag llm-inference api-framework self-hosted crawlers large-language-models web-development developer-tools self-hosted python cli crawler scraper html-to-markdown django scrapy celery sitemap-generation javascript-rendering api-service data-extraction retrieval-augmented-generation docker web-server
7 sources
- readme: https://github.com/watercrawl/WaterCrawl · fetched 2026-08-28 · 47d1f81e03fa
- homepage: https://watercrawl.dev · fetched 2026-08-29 · 196daac95d9c
- site_page: https://docs.watercrawl.dev/clients/python · fetched 2026-08-29 · 9ffa21511a34
- site_page: https://docs.watercrawl.dev/ · fetched 2026-08-29 · 03566b4919e5
- site_page: https://docs.watercrawl.dev/api/documentation · fetched 2026-08-29 · 269d4448523f
- site_page: https://watercrawl.dev/changelog · fetched 2026-08-29 · 35b0c6db7d82
- site_page: https://watercrawl.dev/pricing · fetched 2026-08-29 · bb25210344a9
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| watercrawl/WaterCrawl | main | 82 |
For agents
markdown · JSON · MCP: product_card(name="watercrawl/WaterCrawl")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem