Ross ROSS = Recommend OSS · open-source software intelligence for agents

confident-ai/deepeval

The LLM Evaluation Framework observed · 2026-08-28

github.com/confident-ai/deepeval · homepage · Python · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

95/100

  • Activity 99
  • Release rhythm 99
  • Longevity 79
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 11.5
  • age_days: 1119
  • days_rel: 9
  • days_push: 7
  • n_releases_24m: 27

Full methodology

Adoption not part of the score

17882 stars · 1858 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

DeepEval is an open-source Python framework for evaluating LLM applications with pytest-style unit tests and 50+ research-backed metrics including LLM-as-a-judge, RAG, agent, conversational, safety, and multimodal metrics. It supports end-to-end, trajectory-based, and component-level evals, synthetic dataset generation, and integrations with frameworks like LangChain, LlamaIndex, CrewAI, and OpenAI Agents.

Use cases

  • evaluate rag pipeline faithfulness and hallucination
  • unit test llm outputs in ci/cd with pytest
  • evaluate ai agent trajectories and tool use
  • score chatbot multi-turn conversation quality
  • generate synthetic test datasets for llm edge cases
  • run llm-as-a-judge metrics with custom criteria
  • test llm safety for toxicity and bias
  • trace and evaluate langchain or llamaindex apps

When to choose

  • you need pytest-native LLM evaluation that runs in CI/CD
  • you want ready-made research-backed metrics for RAG, agents, or chatbots
  • you need to evaluate agent execution traces across popular orchestration frameworks
  • you want local-first evaluation with optional cloud dashboards

When to avoid

  • you need a general ML model evaluation library for classical ML tasks
  • you want a hosted-only evaluation service without running code locally
  • your project is not Python-based

Facets

framework · maturity active

testing machine-learning llm-inference rag agent-framework benchmarking data-generation cli large-language-models machine-learning chatbots developer-tools testing python cli cross-platform llm-evaluation llm-as-a-judge pytest evaluation-metrics synthetic-data ci-cd hallucination-detection multimodal-evaluation trajectory-evaluation ai-agents retrieval-augmented-generation

10 sources

Member repositories

RepositoryRoleHealth v2
confident-ai/deepevalmain95

For agents

markdown · JSON · MCP: product_card(name="confident-ai/deepeval")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem