A benchmark is not a scoreboard. It is a contract — a formal statement of what an industry has agreed "good" means. When the dominant benchmarks change, everything downstream of that definition changes with it: procurement criteria, architecture decisions, staffing plans.
OpenAI's GPT-5.6 release in July 2026 is a useful lens on that shift — not because of any single score, but because of which scores the vendor chose to lead with. The headline numbers were not knowledge benchmarks. They were OSWorld, Terminal-Bench, BrowseComp, and a long-horizon professional-workflow evaluation. That editorial choice, replicated across every frontier lab, tells you where the industry believes value now lives: in what models can do, not what they know.
The numbers, classed
Verified, vendor-reported: OpenAI reports GPT-5.6 Sol at 62.6% on OSWorld 2.0 — a harder, long-horizon successor to the original OSWorld benchmark that this release moved to — along with 88.8% on Terminal-Bench 2.1, 90.4% on BrowseComp, and 53.6 on Agents' Last Exam.
Also verified, and more interesting: third-party leaderboards do not agree with these numbers, or with each other. Different trackers show different leaders on the same benchmark, updated at different times.
The disagreement is not sloppiness. It is the central lesson of the computer-use era: a computer-use score is a property of a system, not a model. The number depends on the harness — the scaffolding that renders screens, exposes tools, manages retries, structures memory, and decides when the agent gives up. The same base model under two harnesses can differ by ten points or more. Knowledge benchmarks mostly didn't have this problem; you asked the model a question and it answered. Computer-use benchmarks inherently do.
Corpus concept
Harness Dependence
Measured AI capability is a property of the full system — model, scaffolding, tools, retry policy — not of the model alone. In procurement, require harness disclosure; internally, benchmark candidates inside your own harness, never on public numbers.
Three consequences for buyers
First, vendor-reported agentic scores are best read as demonstrations of what the vendor's own scaffolding achieves — not as portable properties of the model you'll get through an API and wrap in your own architecture.
Second, your architecture is now part of the capability. A mediocre model in an excellent harness will outperform an excellent model in a mediocre one across a wide band of tasks. Your engineering choices are not downstream of model choice — they are peers to it.
Third, the discriminating procurement question is no longer "what did you score?" It's "what harness produced that score, and how does it differ from the deployment configuration you're selling me?"
What to do Monday
Add a harness-disclosure requirement to every AI procurement conversation this quarter: the harness behind every quoted number, its delta from the sold configuration, retry and tool policies, and evaluation conditions. Re-benchmark shortlisted options inside your own harness before deciding — not the vendor's.
And read the benchmark literature whole, not the keynote slice of it. The same field reporting super-human OSWorld scores reports sub-1% completion rates on the hardest professional-workflow tier of Agents' Last Exam. Both numbers are true. The distance between them is where your architecture, your governance, and your competitive advantage live.
This brief reflects vendor-reported figures as of July 2026, identified as such throughout; third-party leaderboard figures vary by harness and update date and were re-verified at publication.