pinchbench/skill resource
PinchBench is a benchmarking system for evaluating LLM models as OpenClaw coding agents. Made with 🦀 by the humans at https://kilo.ai observed · 2026-08-28
Health v2 · maintenance only
72/100
- Activity 90
- Release rhythm 83
- Longevity 14
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 0.0
- age_days: 204
- days_rel: 119
- days_push: 63
- n_releases_24m: 13
Adoption not part of the score
1325 stars · 154 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
PinchBench is a benchmarking system that evaluates LLM models as OpenClaw coding agents using 53 real-world tasks like scheduling, coding, research, and email triage. This repository contains the benchmark skill/task definitions, with results published on a public leaderboard at pinchbench.com.
Use cases
- compare LLM models for coding agent performance
- evaluate which model is best for my AI agent
- benchmark LLM tool usage and multi-step reasoning
- find the best value or cheapest model for agent tasks
- run real-world agent tasks against different models
- measure LLM success rate, speed, and cost
- add custom benchmark tasks for AI agents
When to choose
- You run OpenClaw (or similar) agents and need to pick the best model
- You want real-world task benchmarks rather than synthetic LLM tests
- You need cost/speed/success-rate comparisons across many models via OpenRouter
When to avoid
- You need standard academic benchmarks like SWE-bench or HumanEval
- You don't use OpenClaw or an equivalent agent runtime
- You want to benchmark non-agent capabilities like raw code completion
Facets
dataset · maturity active
benchmarking llm-inference agent-framework testing data-generation large-language-models machine-learning developer-tools testing python cli cross-platform llm-benchmark coding-agents leaderboard openclaw task-suite llm-judge evaluation ai-agents
3 sources
- readme: https://github.com/pinchbench/skill · fetched 2026-08-28 · b5c2e6c273e1
- homepage: https://pinchbench.com · fetched 2026-08-29 · 1e488dfa9617
- site_page: https://pinchbench.com/about · fetched 2026-08-29 · ea2d53367468
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| pinchbench/skill | main | 72 |
For agents
markdown · JSON · MCP: product_card(name="pinchbench/skill")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem