Ross ROSS = Recommend OSS · open-source software intelligence for agents

hjacobs/kubernetes-failure-stories resource

Compilation of public failure/horror stories related to Kubernetes observed · 2026-08-28

github.com/hjacobs/kubernetes-failure-stories · homepage · HTML · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases archived no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2783
  • days_rel: n/a
  • days_push: 2201
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

6224 stars · 330 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

A curated compilation of links to public postmortems and failure stories about Kubernetes outages in production. It is published as the website k8s.af and organized by year, affected components, and impact.

Use cases

  • find real-world kubernetes outage postmortems
  • learn common ways kubernetes clusters fail
  • research production incidents before running k8s
  • find stories about CPU limits, DNS, and CNI failures
  • prepare SRE training material from real incidents
  • avoid repeating known kubernetes pitfalls

When to avoid

  • you need a tool or library rather than reading material
  • you need up-to-date incident reports - the list has not been actively updated since 2023

Facets

learning-resource · maturity maintenance

documentation cloud-computing monitoring kubernetes postmortem sre reliability incident-stories awesome-list devops web-server

2 sources

Member repositories

RepositoryRoleHealth v2
hjacobs/kubernetes-failure-storiesmain10

For agents

markdown · JSON · MCP: product_card(name="hjacobs/kubernetes-failure-stories")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem