Ross ROSS = Recommend OSS · open-source software intelligence for agents

openai/SWELancer-Benchmark resource

This repo contains the dataset and code for the paper "SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?" observed · 2026-08-28

github.com/openai/SWELancer-Benchmark · archived observed · 2026-08-28

Health v2 · maintenance only

10/100

  • Activity 32
  • Release rhythm 35
  • Longevity 40

Flags: no_releases archived no_license

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 561
  • days_rel: n/a
  • days_push: 412
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1431 stars · 135 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

SWE-Lancer is a benchmark dataset and evaluation harness measuring whether frontier LLMs can complete real-world freelance software engineering tasks worth up to $1 million in aggregate. The codebase has been merged into OpenAI's preparedness repository, so this repo now serves as a pointer to the active project.

Use cases

  • evaluate llms on real-world software engineering tasks
  • benchmark coding agents on freelance-style issues
  • compare frontier models on paid software tasks
  • run swe-lancer evaluations
  • research llm coding capability
  • measure ai performance on real github issues

When to choose

  • you need a rigorous benchmark of LLM software engineering ability on real paid tasks
  • you are researching how frontier models handle realistic freelance coding work
  • you want reproducible evaluation harnesses from OpenAI's preparedness project

When to avoid

  • you want an actively developed repo - use openai/preparedness instead
  • you need a general-purpose coding assistant rather than an evaluation benchmark
  • you lack the compute or API budget to run frontier model evaluations

Facets

dataset · maturity maintenance

benchmarking llm-inference machine-learning large-language-models artificial-intelligence developer-tools testing python cross-platform benchmark evaluation llm-evaluation software-engineering freelance-tasks openai swe-benchmark research docker

1 source

Member repositories

RepositoryRoleHealth v2
openai/SWELancer-Benchmarkmain10

For agents

markdown · JSON · MCP: product_card(name="openai/SWELancer-Benchmark")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem