Skip to main content
SandRise logo SandRise
Exploring Next / By company / Swe Bench

Topic

Swe Bench

2 episodes

  1. Ep 577 Jun 30, 2026

    Reward hacking is swamping model intelligence gains · Cursor

    Pippa and Tyler dig into Cursor's claim that coding benchmark gains are being inflated by runtime answer retrieval, not pure model intelligence. They land on the real argument: for historical public-repo evals, the harness is part of the benchmark, because open web and git history can leak the fix and change what the score means.

    EvalsAgentsBenchmarkSwe Bench
  2. Ep 168 Feb 9, 2026

    Group Evolving Agents: Open Ended Self Improvement via Experience Sharing

    Exploring a new paradigm for AI evolution: Group-Evolving Agents. Are they the future or just another research paper?

    AgentsEvalsGroup Evolving AgentsSwe Bench
SandRise logo SandRise Product Studio
Resume LinkedIn GitHub Email

© 2026 SandRise · Built by Nick Sanders

🧠 PM Perspective

Crafting your PM challenge
Analyzing context and generating a thoughtful question...
Your Challenge
0 / 2000
✨

Feedback on Your Answer

⚠️

Say Hi

Feedback, ideas, interesting finds — anything goes.

What's this about?
0 / 2,000

Note received!

Thanks for reaching out. I'll take a look soon.

⚠️

Something went wrong. Please try again.