tech

Separating signal from noise in coding evaluations

Through a detailed audit, we find widespread task issues in SWE-Bench Pro and estimate that ~30% of the tasks are broken.

Separating signal from noise in coding evaluations

TL;DR

  • A detailed audit found approximately 30% of tasks in the SWE-Bench Pro coding benchmark are broken.
  • Issues include overly strict tests, underspecified prompts, low-coverage tests, and misleading prompts.
  • These flaws compromise the benchmark's ability to accurately measure AI coding capabilities and inform safety decisions.
  • The audit used a combination of automated analysis, agent-assisted reviews, and human annotation by experienced software engineers.
  • The findings suggest a retraction of previous recommendations to adopt SWE-Bench Pro.
  • The article emphasizes the difficulty of curating fair benchmarks and the growing utility of AI agents for quality checks.