ianarawjo/ChainForge
An open-source visual programming environment for battle-testing prompts to LLMs. observed · 2026-08-28
Health v2 · maintenance only
66/100
- Activity 86
- Release rhythm 28
- Longevity 89
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 45
- age_days: 1256
- days_rel: 479
- days_push: 84
- n_releases_24m: 4
Adoption not part of the score
3027 stars · 254 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
ChainForge is an open-source visual programming environment for battle-testing prompts to LLMs, built on ReactFlow and Flask. It lets users query multiple LLMs at once, compare responses across prompt permutations and models, and evaluate and visualize results with little to no coding.
Use cases
- compare prompt variations across multiple LLMs
- evaluate LLM response quality with scoring metrics
- visualize prompt evaluation results across models
- test prompt robustness before deploying to production
- run ground-truth evaluations against benchmark datasets
- query OpenAI, Anthropic, Gemini, and Ollama models side by side
- build LLM evaluation flows without writing code
When to choose
- you need to rapidly compare prompts, models, and settings across multiple LLM providers
- you want a visual, low-code interface for LLM evaluation and experimentation
- you need to score and plot LLM responses with custom Python/JavaScript evaluators or LLM-based scorers
- you want to test prompts against local models via Ollama or cloud APIs in one tool
When to avoid
- you need a production LLM pipeline or deployment orchestration rather than experimentation
- you require Safari or other non-Chromium/Firefox browser support
- you want fully automated CI-based eval pipelines with no interactive UI
- you need to avoid executing untrusted code, since evaluator nodes run locally without sandboxing
Facets
application · maturity active
prompt-engineering llm-inference data-visualization machine-learning developer-tools large-language-models artificial-intelligence developer-tools data-visualization python cross-platform browser self-hosted llm-evaluation visual-programming prompt-testing llmops reactflow a-b-testing-prompts model-comparison web-server
10 sources
- readme: https://github.com/ianarawjo/ChainForge · fetched 2026-08-28 · ad0e710fe1fe
- homepage: https://chainforge.ai/docs · fetched 2026-08-29 · f9b4b984a950
- site_page: https://chainforge.ai/docs/vis · fetched 2026-08-29 · 4af11824d20d
- site_page: https://chainforge.ai/docs/getting_started · fetched 2026-08-29 · 7907f20cd5f5
- site_page: https://chainforge.ai/docs/compare_prompts · fetched 2026-08-29 · cc96c7d8bc66
- site_page: https://chainforge.ai/docs/nodes · fetched 2026-08-29 · cfc071743aa9
- site_page: https://chainforge.ai/docs/prompt_templates · fetched 2026-08-29 · 8b46b2c20849
- site_page: https://chainforge.ai/docs/inspection · fetched 2026-08-29 · 1f86f8a5fb8b
- site_page: https://chainforge.ai/docs/evaluation · fetched 2026-08-29 · 28076ab086fa
- site_page: https://chainforge.ai/docs/model_support · fetched 2026-08-29 · a0fe0eb5ef18
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| ianarawjo/ChainForge | main | 66 |
For agents
markdown · JSON · MCP: product_card(name="ianarawjo/ChainForge")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem