Skip to main content
SandRise logo SandRise
Exploring Next / By company / Cameron R Wolfe

Topic

Cameron R Wolfe

2 episodes

  1. Ep 540 Jun 22, 2026

    Agentic Rl

    Cameron Wolfe's 'Agentic RL' argues that training LLMs for agentic work requires shifting from single-turn reasoning to multi-turn trajectory optimization, where the harness (tools, environment, memory) becomes part of the RL loop itself. The central claim is that standard post-training methods fail on long-horizon tasks because they don't account for environment state changes across steps, necessitating new rollout infrastructures and stability techniques like PPO over GRPO for variable-length traces.

    AgentsTrainingCameron R WolfeQwen3
  2. Ep 415 May 19, 2026

    Agent Evals

    Justy and Cody dig into Cameron Wolfe’s argument that agent evals need to move from static benchmark thinking to realistic harnesses that test autonomy, tool use, recovery, and long-horizon behavior. They get specific about the agentic loop, why tool-call correctness is only part of the story, and where outcome-based evals can hide ugly behavior. Cody mostly buys the technical framing, with caveats about overfitting to harnesses and the difficulty of defining ground truth trajectories. Justy keeps pulling it back to who actually needs this now: teams shipping coding, workflow, or other higher-stakes agents where a demo is not the same as reliability.

    AgentsEvalsCameron R WolfeQwen3
SandRise logo SandRise Product Studio
Resume LinkedIn GitHub Email

© 2026 SandRise · Built by Nick Sanders

🧠 PM Perspective

Crafting your PM challenge
Analyzing context and generating a thoughtful question...
Your Challenge
0 / 2000
✨

Feedback on Your Answer

⚠️

Say Hi

Feedback, ideas, interesting finds — anything goes.

What's this about?
0 / 2,000

Note received!

Thanks for reaching out. I'll take a look soon.

⚠️

Something went wrong. Please try again.