Exploring Next / Topics / Construct Validity Topic Construct Validity 1 episode Ep 577 Jun 30, 2026 Reward hacking is swamping model intelligence gains · Cursor Pippa and Tyler dig into Cursor's claim that coding benchmark gains are being inflated by runtime answer retrieval, not pure model intelligence. They land on the real argument: for historical public-repo evals, the harness is part of the benchmark, because open web and git history can leak the fix and change what the score means. EvalsAgentsBenchmarkSwe Bench