Ross ROSS = Recommend OSS · open-source software intelligence for agents

ianarawjo/ChainForge

An open-source visual programming environment for battle-testing prompts to LLMs. observed · 2026-08-28

github.com/ianarawjo/ChainForge · homepage · TypeScript · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

66/100

  • Activity 86
  • Release rhythm 28
  • Longevity 89
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 45
  • age_days: 1256
  • days_rel: 479
  • days_push: 84
  • n_releases_24m: 4

Full methodology

Adoption not part of the score

3027 stars · 254 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

ChainForge is an open-source visual programming environment for battle-testing prompts to LLMs, built on ReactFlow and Flask. It lets users query multiple LLMs at once, compare responses across prompt permutations and models, and evaluate and visualize results with little to no coding.

Use cases

  • compare prompt variations across multiple LLMs
  • evaluate LLM response quality with scoring metrics
  • visualize prompt evaluation results across models
  • test prompt robustness before deploying to production
  • run ground-truth evaluations against benchmark datasets
  • query OpenAI, Anthropic, Gemini, and Ollama models side by side
  • build LLM evaluation flows without writing code

When to choose

  • you need to rapidly compare prompts, models, and settings across multiple LLM providers
  • you want a visual, low-code interface for LLM evaluation and experimentation
  • you need to score and plot LLM responses with custom Python/JavaScript evaluators or LLM-based scorers
  • you want to test prompts against local models via Ollama or cloud APIs in one tool

When to avoid

  • you need a production LLM pipeline or deployment orchestration rather than experimentation
  • you require Safari or other non-Chromium/Firefox browser support
  • you want fully automated CI-based eval pipelines with no interactive UI
  • you need to avoid executing untrusted code, since evaluator nodes run locally without sandboxing

Facets

application · maturity active

prompt-engineering llm-inference data-visualization machine-learning developer-tools large-language-models artificial-intelligence developer-tools data-visualization python cross-platform browser self-hosted llm-evaluation visual-programming prompt-testing llmops reactflow a-b-testing-prompts model-comparison web-server

10 sources

Member repositories

RepositoryRoleHealth v2
ianarawjo/ChainForgemain66

For agents

markdown · JSON · MCP: product_card(name="ianarawjo/ChainForge")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem