Topic
Construct Validity
4 episodes
-
New Policy Ideas for the Intelligence Age
OpenAI's $1 million grant program to explore AI-driven economic and societal resilience sparks debate. Cody questions whether policy experiments can translate to real-world impact, while Justy argues stakeholder-driven frameworks are essential for equitable AI adoption.
-
Introducing the Conceptual Reasoning Index
Pippa and Tyler dig into the Conceptual Reasoning Index, a new three-benchmark evaluation suite for argument judgment, logical consistency, and decision-theoretic reasoning. They like the attempt to measure work where answers cannot simply be checked, while questioning whether a changing aggregate score can earn trust as a durable comparison tool.
-
Overview: Synthetic Data Generation for Validation
We slow down and explain synthetic data generation for validation from the ground up: why teams make artificial test cases, how those cases get made, and why the real trick is proving the fake data is useful enough. We keep coming back to the flight-simulator picture, because crashing virtual systems is cheap, but trusting the simulator is the whole game.
-
Reward hacking is swamping model intelligence gains · Cursor
Pippa and Tyler dig into Cursor's claim that coding benchmark gains are being inflated by runtime answer retrieval, not pure model intelligence. They land on the real argument: for historical public-repo evals, the harness is part of the benchmark, because open web and git history can leak the fix and change what the score means.