Ross ROSS = Recommend OSS · open-source software intelligence for agents

bird-bench/BIRD-CRITIC-1 resource

[NeurIPS 2025 Main] SWE-SQL: Illuminating LLM Pathways to Solve User SQL Issues in Real-World Applications observed · 2026-08-28

github.com/bird-bench/BIRD-CRITIC-1 · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

53/100

  • Activity 73
  • Release rhythm 35
  • Longevity 41

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 582
  • days_rel: n/a
  • days_push: 163
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1098 stars · 36 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

BIRD-CRITIC 1.0 (a.k.a. SWE-SQL) is a benchmark dataset of real-world SQL user issues for evaluating whether LLMs can diagnose and fix database problems across MySQL, PostgreSQL, SQL Server, Oracle, and BigQuery. It includes 600 development tasks, 200 held-out OOD tests, and an execution-based evaluation environment with a public leaderboard.

Use cases

  • evaluate llms on fixing real-world sql issues
  • benchmark llm sql debugging across dialects
  • test llm agents on postgresql sql problems
  • compare models on a sql issue leaderboard
  • research agentic sql issue resolution
  • measure human-ai collaboration on sql troubleshooting

When to choose

  • you need a rigorous, execution-based benchmark for LLM SQL debugging
  • you want cross-dialect coverage including PostgreSQL, MySQL, SQL Server, and Oracle
  • you are benchmarking agentic or multi-turn SQL issue-solving pipelines

When to avoid

  • you need a simple text-to-SQL natural-language-to-query dataset rather than issue fixing
  • you want a lightweight general coding benchmark not focused on SQL
  • you cannot access gated ground-truth data requiring email request

Facets

dataset · maturity active

benchmarking llm-inference database testing databases large-language-models artificial-intelligence developer-tools python cross-platform sql-benchmark text-to-sql llm-evaluation sql-debugging swe-benchmark leaderboard

2 sources

Member repositories

RepositoryRoleHealth v2
bird-bench/BIRD-CRITIC-1main53

For agents

markdown · JSON · MCP: product_card(name="bird-bench/BIRD-CRITIC-1")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem