Introducing GeneBench-Pro
modelsresearch
read at source ↗ openai.com
Introducing GeneBench-Pro
Source: OpenAI Date: 2026-06-30 URL: https://openai.com/index/introducing-genebench-pro
Summary
OpenAI released GeneBench-Pro, a 129-problem benchmark testing whether agents can navigate messy, noisy real-world biological data (genomics, quantitative biology, translational medicine) and make the judgment calls — handling measurement error, confounding, QC failures, model selection — that actual computational research requires, with problems estimated at 20-40 human-expert-hours each. GPT-5.6 Sol scored 28.7% (31.5% in Pro mode), signaling the benchmark is intentionally near-unsolved.
Implications
- Feeds the two capability clocks (open vs closed frontier) — a closed-frontier lab publishing a benchmark its own flagship model fails at ~70% of the time is a deliberate frontier-marking move, distinct from capability announcements; watch whether any open model attempts GeneBench-Pro at all, since a large closed/open gap here would be a genuine differentiator (unlike coding benchmarks where open models are closing fast).
- Feeds agent-layer convergence — the benchmark targets multi-stage agentic workflows with judgment calls under noise, the same shape of problem (long-horizon, ambiguous, verification-hard) driving harness-level reliability fixes elsewhere this week.