langwatch/langwatch
The platform for LLM evaluations and AI agent testing observed · 2026-08-28
Health v2 · maintenance only
90/100
- Activity 99
- Release rhythm 87
- Longevity 77
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 0.0
- age_days: 1089
- days_rel: 7
- days_push: 7
- n_releases_24m: 219
Adoption not part of the score
3511 stars · 355 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
LangWatch is an open-source LLMOps platform for LLM evaluation, AI agent testing, tracing, and production observability. It combines agent simulations, dataset-driven experiments, online evaluation monitors, prompt management, and an OpenAI-compatible AI gateway, deployable as a cloud service or self-hosted via Helm.
Use cases
- evaluate llm outputs for quality and safety
- run agent simulations before release
- monitor production llm traffic for regressions
- trace and debug ai agent behavior
- catch prompt regressions in ci/cd
- red team agents for jailbreaks and unsafe tool calls
- manage prompts and optimize models across providers
- self-host an llm observability platform
When to choose
- you need end-to-end evaluation, tracing, and prompt management in one platform
- you want simulation-based agent testing with simulated users and LLM judges
- you need OpenTelemetry-native, provider-agnostic instrumentation
- you require self-hosting for GDPR/SOC 2 or air-gapped environments
When to avoid
- you only need simple token/cost logging without evaluation workflows
- you want a fully free managed service with no usage-based pricing
- you need a lightweight library rather than a full platform deployment
Facets
service · maturity active
monitoring tracing testing benchmarking analytics prompt-engineering rag llm-inference api-gateway sdk large-language-models machine-learning developer-tools monitoring testing self-hosted self-hosted python cloud llm-evaluation llmops agent-testing agent-simulation opentelemetry red-teaming guardrails prompt-optimization dspy open-core ai-agents docker kubernetes web-server nodejs
10 sources
- readme: https://github.com/langwatch/langwatch · fetched 2026-08-28 · 1b448fc450da
- homepage: https://langwatch.ai · fetched 2026-08-29 · abecdc11f9e9
- site_page: https://langwatch.ai/docs/skills/directory · fetched 2026-08-29 · 6b5b8d2c7edc
- site_page: https://langwatch.ai/docs/evaluations/evaluators/overview · fetched 2026-08-29 · 3852856b4b34
- site_page: https://langwatch.ai/docs/evaluations/experiments/overview · fetched 2026-08-29 · 16ffb7de427f
- site_page: https://docs.langwatch.ai · fetched 2026-08-29 · cdbdb4f37f38
- site_page: https://langwatch.ai/docs/self-hosting/overview · fetched 2026-08-29 · f5d5d1863575
- site_page: https://langwatch.ai/docs/evaluations/online-evaluation/overview · fetched 2026-08-29 · 1ba0669f9892
- registry_npm: https://registry.npmjs.org/langwatch · fetched 2026-08-29 · e60b6314db43
- site_page: https://langwatch.ai/scenario/agent-integration · fetched 2026-08-29 · c67280ce27b0
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| langwatch/langwatch | main | 90 |
For agents
markdown · JSON · MCP: product_card(name="langwatch/langwatch")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem