Ross ROSS = Recommend OSS · open-source software intelligence for agents

bird-bench/BIRD-Interact resource

[ICLR 2026 Oral] BIRD-INTERACT: Re-imagines Text-to-SQL evaluation via lens of dynamic interactions. observed · 2026-08-28

github.com/bird-bench/BIRD-Interact · homepage · Python · MIT (permissive) observed · 2026-08-28

Health v2 · maintenance only

52/100

  • Activity 74
  • Release rhythm 35
  • Longevity 33

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 469
  • days_rel: n/a
  • days_push: 157
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

1011 stars · 25 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

BIRD-INTERACT is an interactive Text-to-SQL benchmark that evaluates LLMs through dynamic multi-turn interactions with a simulated user, a hierarchical knowledge base, and database documentation. It provides 600 annotated tasks spanning BI and CRUD operations in passive conversational and active agentic modes, each guarded by executable test cases.

Use cases

  • evaluate LLMs on interactive text-to-sql tasks
  • benchmark agentic database assistants
  • test conversational SQL generation with clarifying questions
  • measure model performance on CRUD and BI database tasks
  • compare reasoning models on multi-turn database interaction
  • research dynamic evaluation for text-to-sql

When to choose

  • you need a rigorous, executable-test-guarded benchmark for interactive Text-to-SQL
  • you want to evaluate agents that must ask clarifying questions or act proactively
  • you are benchmarking production-ready database assistant capabilities

When to avoid

  • you need a simple static text-to-SQL dataset without interaction
  • you need a lightweight single-turn SQL generation benchmark
  • your focus is non-SQL code generation evaluation

Facets

dataset · maturity active

benchmarking llm-inference agent-framework database nlp databases large-language-models artificial-intelligence python cli text-to-sql benchmark interactive-evaluation user-simulator llm-evaluation sql ai-agents natural-language-processing

2 sources

Member repositories

RepositoryRoleHealth v2
bird-bench/BIRD-Interactmain52

For agents

markdown · JSON · MCP: product_card(name="bird-bench/BIRD-Interact")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem