# uptrain-ai/uptrain

UpTrain is an open-source unified platform to evaluate and improve Generative AI applications. We provide grades for 20+ preconfigured checks (covering language, code, embedding use-cases), perform root cause analysis on failure cases and give insights on how to resolve them.

Repository: https://github.com/uptrain-ai/uptrain
Canonical: https://ross.abutalabs.com/products/uptrain
Homepage: https://uptrain.ai/
Language: Python
License: Apache-2.0
License Family: permissive
Topics: machine-learning, experimentation, llm-prompting, llm-test, llmops, monitoring, prompt-engineering, autoevaluation, evaluation, llm-eval, hallucination-detection, jailbreak-detection, openai-evals, root-cause-analysis
Last push: 2024-08-18T13:30:44+00:00

## Health v2 (maintenance only)
Score: 23/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 0, release rhythm 8, longevity 99
- inputs: {"age_days": 1395, "days_push": 745, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2362, forks 205 (observed 2026-08-28T04:06:40.981448+00:00)

## What it is
UpTrain is an open-source Python platform for evaluating and improving LLM applications, offering 20+ preconfigured evaluation checks covering language, code, and embedding use cases. It performs root cause analysis on failure cases, supports automated regression testing and prompt versioning, and includes locally run dashboards for visualizing results.

## Use cases
- evaluate llm responses for factual accuracy and hallucinations
- test rag pipeline retrieval quality
- run regression tests on prompt changes
- find root causes of llm application failures
- score chatbot response quality and tonality
- detect jailbreak attempts in llm outputs
- build custom llm evaluation metrics

## When to choose
- you need automated evaluation of LLM application outputs with prebuilt metrics
- you want root cause analysis and failure pattern detection for RAG or chatbot pipelines
- you need regression testing and prompt versioning for LLM apps
- you prefer an open-source, self-hostable LLMOps evaluation tool

## When to avoid
- you need full observability/tracing of production LLM traffic rather than evaluation
- you want a managed SaaS-only evaluation service
- your stack is not Python-based
- you need evaluation of non-LLM machine learning models

## Facets
- artifact type: library
- maturity: active
- function: testing, monitoring, benchmarking, llm-inference, prompt-engineering, rag, analytics
- domain: large-language-models, machine-learning, developer-tools, artificial-intelligence, testing
- platform: python, self-hosted
- tags: llm-evaluation, llmops, hallucination-detection, root-cause-analysis, rag-evaluation, regression-testing, observability, docker

## Member repositories
- uptrain-ai/uptrain (main) score 23

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:40.981448+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:36:05.154676+00:00, confidence not recorded.
  - readme: https://github.com/uptrain-ai/uptrain (fetched 2026-08-28T04:06:40.981448+00:00, sha 90ad016523af)
  - homepage: https://uptrain.ai/ (fetched 2026-08-29T10:16:38.711327+00:00, sha 74863bf39c8e)
  - site_page: https://docs.uptrain.ai/ (fetched 2026-08-29T10:16:38.720441+00:00, sha 4dcf57ea43c7)
  - registry_pypi: https://pypi.org/pypi/uptrain/json (fetched 2026-08-29T10:16:38.722246+00:00, sha eaf02a130911)
- Data as of 2026-08-30T08:39:29.467469+00:00.
