# HolmesGPT/holmesgpt

SRE Agent - CNCF Sandbox Project

Repository: https://github.com/HolmesGPT/holmesgpt
Canonical: https://ross.abutalabs.com/products/holmesgpt
Homepage: https://holmesgpt.dev/
Language: Python
License: Apache-2.0
License Family: permissive
Topics: aiops, kubernetes, llm, llm-agent, llm-framework, llms, monitoring, observability, prometheus, chatbot, chatops, devops, incident, incident-management, incident-response, sre, jira, slack, devops-tools, site-reliability-engineering
Last push: 2026-08-26T07:26:34+00:00

## Health v2 (maintenance only)
Score: 87/100 (v2, computed 2026-09-03T02:39:23.370411+00:00)
- activity 99, release rhythm 87, longevity 58
- inputs: {"age_days": 825, "days_push": 7, "days_rel": 7, "gap_med": 6, "n_releases_24m": 74}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3148, forks 454 (observed 2026-08-28T04:07:46.021530+00:00)

## What it is
HolmesGPT is an open-source AI agent (CNCF Sandbox project) that investigates production incidents and finds root causes by querying live observability data from sources like Kubernetes, Prometheus, Grafana, and Datadog. It supports an agentic loop, background operator mode with scheduled health checks, and integrations with Slack, Jira, PagerDuty, and any LLM provider.

## Use cases
- investigate kubernetes pod crashloop root cause
- automate incident triage with an LLM agent
- get slack alerts with AI root cause analysis
- run scheduled health checks on my services
- troubleshoot prometheus alerts automatically
- find root cause of production incidents with AI
- connect datadog and grafana to an AI SRE agent

## When to choose
- you run Kubernetes or cloud infrastructure and want automated incident investigation
- you want an AI agent that queries live observability data instead of static docs
- you need alert enrichment and root cause analysis piped into Slack, Jira, or PagerDuty
- you want background 24/7 problem detection with operator mode

## When to avoid
- you need a lightweight metrics scraper or dashboard rather than an LLM agent
- you cannot send infrastructure data to an LLM provider due to compliance constraints
- you want a general-purpose chatbot unrelated to ops or observability

## Facets
- artifact type: application
- maturity: active
- function: agent-framework, llm-inference, monitoring, alerting, chatbot, rag
- domain: monitoring, artificial-intelligence, large-language-models, self-hosted
- platform: python, self-hosted, cli
- tags: sre, incident-response, root-cause-analysis, observability, aiops, chatops, prometheus, kubernetes, slack, jira, cncf, devops, ai-agents, docker

## Member repositories
- HolmesGPT/holmesgpt (main) score 87

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:46.021530+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:45:35.193071+00:00, confidence not recorded.
  - readme: https://github.com/HolmesGPT/holmesgpt (fetched 2026-08-28T04:07:46.021530+00:00, sha 8baee033d93e)
  - homepage: https://holmesgpt.dev/ (fetched 2026-08-29T09:40:31.610000+00:00, sha 36c6c3c2e4f9)
- Data as of 2026-08-30T08:39:29.467469+00:00.
