# datacurve-ai/deep-swe

Measuring frontier coding agents on original, long-horizon engineering tasks

Repository: https://github.com/datacurve-ai/deep-swe
Canonical: https://ross.abutalabs.com/products/deep-swe
Homepage: https://deepswe.datacurve.ai/
Language: Python
License: Apache-2.0
License Family: permissive
Last push: 2026-08-06T20:05:06+00:00

## Health v2 (maintenance only)
Score: 57/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 96, release rhythm 35, longevity 7
- inputs: {"age_days": 110, "days_push": 27, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: no_releases, young
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1499, forks 99 (observed 2026-08-28T04:04:54.059958+00:00)

## What it is
DeepSWE is a benchmark of 113 original, long-horizon software engineering tasks drawn from active open-source repositories, used to measure frontier coding agents across TypeScript, Go, Python, JavaScript, and Rust. It ships with isolated Docker environments, program-based verifiers, and a public leaderboard of model results.

## Use cases
- evaluate coding agents on realistic long-horizon engineering tasks
- compare frontier LLMs on software engineering benchmarks
- benchmark a new model against Claude, GPT, Gemini, and DeepSeek results
- run sandboxed agent evaluations with network allowlists
- measure pass@1 and cost efficiency of coding agents

## When to choose
- you need a rigorous, verifiable benchmark for coding agents
- you want up-to-date leaderboard comparisons across many LLMs
- you need isolated, reproducible eval environments with held-out tests

## When to avoid
- you need a lightweight quick eval rather than full Docker-based runs
- you want short single-function coding tasks instead of long-horizon work
- you cannot run Docker or lack API access to frontier models

## Facets
- artifact type: dataset
- maturity: active
- function: benchmarking, testing, agent-framework, developer-tools
- domain: artificial-intelligence, large-language-models, developer-tools, testing
- platform: python, cross-platform
- tags: coding-agents, benchmark, llm-evaluation, software-engineering-tasks, leaderboard, harbor-format, ai-agents, docker

## Member repositories
- datacurve-ai/deep-swe (main) score 57

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:54.059958+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:33:02.328193+00:00, confidence not recorded.
  - readme: https://github.com/datacurve-ai/deep-swe (fetched 2026-08-28T04:04:54.059958+00:00, sha 6bc496ee389e)
  - homepage: https://deepswe.datacurve.ai/ (fetched 2026-08-29T11:38:09.597746+00:00, sha 6a65a7f71b2a)
  - site_page: https://deepswe.datacurve.ai/changelog (fetched 2026-08-29T11:38:09.607286+00:00, sha 72cd8df7390e)
- Data as of 2026-08-30T08:39:29.467469+00:00.
