Ross ROSS = Recommend OSS · open-source software intelligence for agents

SWE-bench/SWE-bench resource

SWE-bench: Can Language Models Resolve Real-world Github Issues? observed · 2026-08-28

github.com/SWE-bench/SWE-bench · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

72/100

  • Activity 98
  • Release rhythm 35
  • Longevity 76

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 1065
  • days_rel: n/a
  • days_push: 15
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

5719 stars · 953 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

SWE-bench is a benchmark and evaluation harness that tests whether large language models can resolve real-world GitHub issues by generating patches against actual codebases. It ships datasets (SWE-bench, Lite, Verified, Multimodal), a Docker-based reproducible evaluation harness, and official leaderboards.

Use cases

  • evaluate llm on real github issues
  • benchmark coding agents on software engineering tasks
  • measure how well a model generates bug-fix patches
  • compare coding assistants on a public leaderboard
  • run reproducible docker-based code evaluations
  • find a dataset of real software engineering problems for llm research

When to choose

  • you need a standardized, reproducible benchmark for LLM software-engineering ability
  • you want to compare your model or agent against published state-of-the-art results
  • you need real-world issue/patch data rather than synthetic coding tasks

When to avoid

  • you need a general coding benchmark for simple algorithmic problems (e.g., HumanEval-style)
  • you cannot run Docker or containerized evaluations
  • you want a lightweight unit-testing framework rather than an LLM evaluation suite

Facets

dataset · maturity active

benchmarking testing machine-learning llm-inference agent-framework large-language-models developer-tools machine-learning testing python llm-benchmark code-generation github-issues evaluation-harness leaderboard swe-agent patch-generation software-engineering docker linux macos

3 sources

Member repositories

RepositoryRoleHealth v2
SWE-bench/SWE-benchmain72

For agents

markdown · JSON · MCP: product_card(name="SWE-bench/SWE-bench")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem