Ross ROSS = Recommend OSS · open-source software intelligence for agents

hamelsmu/evals-skills

Skills for AI Evals to compliment the course: AI Evals For Engineers & PMs observed · 2026-08-28

github.com/hamelsmu/evals-skills · homepage · MIT (permissive) · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 98
  • Release rhythm 35
  • Longevity 13

Flags: no_releases archived

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 185
  • days_rel: n/a
  • days_push: 17
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1664 stars · 169 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A collection of skills (prompts/instructions) that guide AI coding agents like Claude Code through building and auditing LLM evaluation pipelines. It includes skills for eval audits, error analysis, synthetic data generation, LLM-as-Judge prompt design, evaluator calibration, RAG evaluation, and annotation interfaces.

Use cases

  • audit my LLM eval pipeline for common mistakes
  • categorize failures from LLM traces with error analysis
  • generate diverse synthetic test inputs for evals
  • design an LLM-as-Judge prompt for subjective quality criteria
  • calibrate an LLM judge against human labels
  • evaluate retrieval and generation quality in a RAG pipeline
  • build an annotation interface for human trace review

When to choose

  • you use Claude Code or a skills-compatible coding agent and want expert guidance on LLM evals
  • you are building or improving an eval pipeline and want to catch common mistakes
  • you need structured workflows for error analysis, judge calibration, or RAG evaluation
  • you are a student of the AI Evals course applying its methodology

When to avoid

  • you want a standalone eval framework or library to run in CI without a coding agent
  • you need a general-purpose testing tool unrelated to LLM evaluation
  • you don't use an AI coding agent that supports the skills/plugin format
  • you want the actively maintained version - this repo is deprecated in favor of ai-evals-course/evals-skills

Facets

plugin · maturity maintenance

testing llm-inference agent-framework prompt-engineering rag developer-tools large-language-models artificial-intelligence developer-tools testing education cli cross-platform llm-evals claude-code-plugin skills llm-as-judge error-analysis synthetic-data rag-evaluation annotation-interface deprecated-repo nodejs

4 sources

Member repositories

RepositoryRoleHealth v2
hamelsmu/evals-skillsmain10

For agents

markdown · JSON · MCP: product_card(name="hamelsmu/evals-skills")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem