# langwatch/langwatch

The platform for LLM evaluations and AI agent testing

Repository: https://github.com/langwatch/langwatch
Canonical: https://ross.abutalabs.com/products/langwatch
Homepage: https://langwatch.ai
Language: TypeScript
License: Apache-2.0
License Family: permissive
Topics: ai, analytics, datasets, evaluation, gpt, llm, observability, openai, prompt-engineering, dspy, llmops, low-code, llm-ops
Last push: 2026-08-26T22:40:44+00:00

## Health v2 (maintenance only)
Score: 90/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 99, release rhythm 87, longevity 77
- inputs: {"age_days": 1089, "days_push": 7, "days_rel": 7, "gap_med": 0.0, "n_releases_24m": 219}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3511, forks 355 (observed 2026-08-28T04:08:07.702587+00:00)

## What it is
LangWatch is an open-source LLMOps platform for LLM evaluation, AI agent testing, tracing, and production observability. It combines agent simulations, dataset-driven experiments, online evaluation monitors, prompt management, and an OpenAI-compatible AI gateway, deployable as a cloud service or self-hosted via Helm.

## Use cases
- evaluate llm outputs for quality and safety
- run agent simulations before release
- monitor production llm traffic for regressions
- trace and debug ai agent behavior
- catch prompt regressions in ci/cd
- red team agents for jailbreaks and unsafe tool calls
- manage prompts and optimize models across providers
- self-host an llm observability platform

## When to choose
- you need end-to-end evaluation, tracing, and prompt management in one platform
- you want simulation-based agent testing with simulated users and LLM judges
- you need OpenTelemetry-native, provider-agnostic instrumentation
- you require self-hosting for GDPR/SOC 2 or air-gapped environments

## When to avoid
- you only need simple token/cost logging without evaluation workflows
- you want a fully free managed service with no usage-based pricing
- you need a lightweight library rather than a full platform deployment

## Facets
- artifact type: service
- maturity: active
- function: monitoring, tracing, testing, benchmarking, analytics, prompt-engineering, rag, llm-inference, api-gateway, sdk
- domain: large-language-models, machine-learning, developer-tools, monitoring, testing, self-hosted
- platform: self-hosted, python, cloud
- tags: llm-evaluation, llmops, agent-testing, agent-simulation, opentelemetry, red-teaming, guardrails, prompt-optimization, dspy, open-core, ai-agents, docker, kubernetes, web-server, nodejs

## Member repositories
- langwatch/langwatch (main) score 90

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:08:07.702587+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:35:43.313611+00:00, confidence not recorded.
  - readme: https://github.com/langwatch/langwatch (fetched 2026-08-28T04:08:07.702587+00:00, sha 1b448fc450da)
  - homepage: https://langwatch.ai (fetched 2026-08-29T09:30:01.482298+00:00, sha abecdc11f9e9)
  - site_page: https://langwatch.ai/docs/skills/directory (fetched 2026-08-29T09:30:01.498337+00:00, sha 6b5b8d2c7edc)
  - site_page: https://langwatch.ai/docs/evaluations/evaluators/overview (fetched 2026-08-29T09:30:01.508541+00:00, sha 3852856b4b34)
  - site_page: https://langwatch.ai/docs/evaluations/experiments/overview (fetched 2026-08-29T09:30:01.510319+00:00, sha 16ffb7de427f)
  - site_page: https://docs.langwatch.ai (fetched 2026-08-29T09:30:01.491533+00:00, sha cdbdb4f37f38)
  - site_page: https://langwatch.ai/docs/self-hosting/overview (fetched 2026-08-29T09:30:01.496104+00:00, sha f5d5d1863575)
  - site_page: https://langwatch.ai/docs/evaluations/online-evaluation/overview (fetched 2026-08-29T09:30:01.511952+00:00, sha 1ba0669f9892)
  - registry_npm: https://registry.npmjs.org/langwatch (fetched 2026-08-29T09:30:01.513519+00:00, sha e60b6314db43)
  - site_page: https://langwatch.ai/scenario/agent-integration (fetched 2026-08-29T09:30:01.506248+00:00, sha c67280ce27b0)
- Data as of 2026-08-30T08:39:29.467469+00:00.
