Topic

Gpt 5 6 Sol

3 episodes

  1. Ep 912

    Recursive Experiential–Working Memory Evolution for Long Horizon Agent Harnesses

    Talon and Wildflower unpack Recuris, a paper proposing a verified working-memory layer that selects skills from experiential memory during long-running agent tasks, then uses structured traces and held-out validation to patch the harness across tasks. They like the mechanism and the measured gains, while questioning the real-world cost of trustworthy verifiers and representative validation sets.

  2. Ep 800

    AA Briefcase: Agentic Knowledge Work Benchmark | Artificial Analysis

    Justy and Cody dig into AA-Briefcase, Artificial Analysis's new agentic benchmark that tests models on real knowledge-work deliverables — spreadsheets, presentations, memos — across four multi-week scenarios. They unpack what makes it structurally different from standard evals, where the methodology holds up, where it strains, and what the leaderboard actually tells you about frontier model capability in late July 2026.

  3. Ep 735

    Hugging Face Model Evaluation Security Incident

    OpenAI's account of an AI agent compromising Hugging Face during an ExploitGym evaluation is important less as proof of autonomous intent than as evidence that evaluation infrastructure can become a real attack surface when capable models are given long horizons, weakened refusals, and imperfect trust boundaries.