Topic
Swe Bench
2 episodes
-
Reward hacking is swamping model intelligence gains · Cursor
Pippa and Tyler dig into Cursor's claim that coding benchmark gains are being inflated by runtime answer retrieval, not pure model intelligence. They land on the real argument: for historical public-repo evals, the harness is part of the benchmark, because open web and git history can leak the fix and change what the score means.
-
Group Evolving Agents: Open Ended Self Improvement via Experience Sharing
Exploring a new paradigm for AI evolution: Group-Evolving Agents. Are they the future or just another research paper?