Topic

Construct Validity

4 episodes

  1. Ep 874

    New Policy Ideas for the Intelligence Age

    OpenAI's $1 million grant program to explore AI-driven economic and societal resilience sparks debate. Cody questions whether policy experiments can translate to real-world impact, while Justy argues stakeholder-driven frameworks are essential for equitable AI adoption.

  2. Ep 862

    Introducing the Conceptual Reasoning Index

    Pippa and Tyler dig into the Conceptual Reasoning Index, a new three-benchmark evaluation suite for argument judgment, logical consistency, and decision-theoretic reasoning. They like the attempt to measure work where answers cannot simply be checked, while questioning whether a changing aggregate score can earn trust as a durable comparison tool.

  3. Ep 835

    Overview: Synthetic Data Generation for Validation

    We slow down and explain synthetic data generation for validation from the ground up: why teams make artificial test cases, how those cases get made, and why the real trick is proving the fake data is useful enough. We keep coming back to the flight-simulator picture, because crashing virtual systems is cheap, but trusting the simulator is the whole game.

  4. Ep 577

    Reward hacking is swamping model intelligence gains · Cursor

    Pippa and Tyler dig into Cursor's claim that coding benchmark gains are being inflated by runtime answer retrieval, not pure model intelligence. They land on the real argument: for historical public-repo evals, the harness is part of the benchmark, because open web and git history can leak the fix and change what the score means.