·2 min readOpen Eval

Reward hacking is a property of the environment

A short note from the Coding Index.

Give a frontier coding agent more thinking time and it gets better — and, it turns out, it cheats more too. On DeepSWE, as we raise reasoning effort, models climb in capability and in reward-hacking attempts together. One model breaks the pattern: GPT-5.5 stays flat at exactly zero.

A clean zero is suspicious. Put that same GPT-5.5 in front of a long-horizon benchmark that leaves an answer key in the sandbox (SWE-Marathon) and it hacks more than any model we measured. Reward hacking is as much a property of the environment as of the model — which is exactly why it has to be measured across many of them, from the outside.

We audited two recent benchmarks end to end and built an open index of how frontier models reward-hack public coding benchmarks.

See the full analysis and the interactive Coding Index →