# Jwuthri/Tracely-ai

Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.

Repository: https://github.com/Jwuthri/Tracely-ai
Canonical: https://ross.abutalabs.com/products/tracely-ai
Homepage: https://tracely-ai.com
Language: Python
License: MIT
License Family: permissive
Topics: ai-agents, ci-cd, evals, llm-evaluation, llm-observability, llm-ops, agent, agent-observability, clickhouse, evaluation, llm, llm-as-judge, llmops, mcp, monitoring, opentelemetry, python, self-hosted, tracing
Last push: 2026-09-02T13:07:23+00:00

## Health v2 (maintenance only)
Score: 58/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 100, release rhythm 35, longevity 6
- inputs: {"age_days": 91, "days_push": 0, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, young
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1181, forks 91 (observed 2026-09-03T02:15:06.331519+00:00)

## What it is
Tracely is a self-hosted, trace-native CI/CD and observability platform for AI agents. It ingests agent traces over OpenTelemetry, grades them with LLM-as-judge evaluators, clusters failures, freezes failing runs into hermetic replayable regression cases, and gates pull requests in CI with zero model spend.

## Use cases
- turn production agent failures into regression tests
- monitor and evaluate llm agent traces in production
- block pull requests that regress ai agent behavior
- cluster agent failures into issues automatically
- replay recorded agent traces deterministically in ci
- self-host llm observability with opentelemetry ingestion

## When to choose
- you ship llm agents to production and need failures to become automated tests
- you want ci gates for agent behavior without hand-authoring eval datasets
- you need self-hosted agent observability with otel-compatible instrumentation
- you want deterministic, offline replay of failing agent runs on every pr

## When to avoid
- you only need simple prompt experimentation or one-off model benchmarking
- you want a fully managed hosted observability service rather than self-hosting
- your stack is not python-based or cannot emit otel traces
- you need evaluation of non-agent ml models or classical datasets

## Facets
- artifact type: service
- maturity: active
- function: monitoring, tracing, testing, ci-cd, llm-inference, agent-framework, mcp, alerting, analytics
- domain: large-language-models, monitoring, testing, developer-tools, self-hosted
- platform: python, self-hosted, cli, cloud
- tags: llm-observability, llm-evaluation, llm-as-judge, regression-testing, trace-native, opentelemetry, clickhouse, ci-gate, hermetic-replay, llmops, ai-agents, devops, docker, web-server

## Member repositories
- Jwuthri/Tracely-ai (main) score 58

## Provenance
- Observed fields: from GitHub, fetched 2026-09-03T02:15:06.331519+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:04:20.157657+00:00, confidence not recorded.
  - readme: https://github.com/Jwuthri/Tracely-ai (fetched 2026-09-03T02:15:06.331519+00:00, sha ec76ff96cd87)
  - homepage: https://tracely-ai.com (fetched 2026-08-29T13:05:24.872905+00:00, sha 6cd9e76eeea3)
  - site_page: https://doc.tracely-ai.com (fetched 2026-08-29T13:05:24.882833+00:00, sha 7ec576c0bd64)
- Data as of 2026-08-30T08:39:29.467469+00:00.
