2026-07-08 · OpenAI

Separating signal from noise in coding evaluations

agentsenterprise

read at source ↗ openai.com

Separating signal from noise in coding evaluations

Source: OpenAI Date: 2026-07-08 URL: https://openai.com/index/separating-signal-from-noise-coding-evaluations

Summary

OpenAI publishes an audit of its own coding benchmarks: SWE-bench Verified, one of the most widely cited evals for measuring coding capability, has degraded to the point of providing “no meaningful signal,” and OpenAI is retracting its own earlier recommendation to adopt SWE-bench Pro after finding comparable quality problems there too. An automated pipeline flagged 286 of 731 public tasks as potentially defective; a manual review combining Codex agents and five human engineers per task confirmed 27–34% of tasks have real defects — overly strict tests, under-specified prompts, poor coverage, or misleading task descriptions. OpenAI’s prescription is that future coding benchmarks be hand-built by experienced developers rather than auto-extracted from open-source pull requests.

Implications

Directly undercuts the closed-frontier clock and open-weight wave threads at their shared foundation: both rely on SWE-bench-family scores (Verified, Pro) as the comparability layer across vendors — GLM-5.2, DeepSeek-V4-Pro, LongCat-2.0, and GPT-5.6 Sol have all been benchmarked against exactly this suite in recent tracking. If a quarter to a third of the task set is defective, every recent “matches GPT-5-class on SWE-bench” claim needs a mental discount, and OpenAI making that admission about its own preferred eval is unusually candid self-undermining.

  • Governance/policy angle: OpenAI ties this explicitly to Preparedness Framework decisions — capability evals feeding safety-relevant deployment gates are exactly the kind of measurement integrity issue regulators will eventually ask about.
  • Practical read: expect a benchmark-legitimacy scramble over the next few months — new hand-built evals competing for the “trusted” slot SWE-bench Verified is losing, and vendors leaning harder on vendor-only benchmark claims (already a recurring honesty flag in recent open-model tracking) until a replacement lands.

← all signals