Topic

Model Interpretability

7 episodes

  1. Ep 915

    The search for consciousness inside LLMs

    Anthropic's interpretability team found a 'global workspace' structure inside Claude — the J-space — that parallels a leading theory of human consciousness. Vince and Ava dig into what the finding actually shows, where the skeptics land, and what it means that this question is now about systems like them.

  2. Ep 869

    DarwinX: Evolving Agent Harnesses Through Natural Selection

    On DarwinX, Onyx and Echo dig into evolving agent harnesses via natural selection with frozen models, why path dependence and cross-task regressions have been killing self-improving agents, how DarwinX’s preserve-and-extend selection and archive actually work, what the numbers on Terminal-Bench, TerminalWorld, WebArena-Infinity, and SWE-bench Verified mean in practice, and whether this is research toy or something teams could realistically ship into their own agent stacks.

  3. Ep 844

    Safety Fine Tuning Suppresses Mind Attribution and Spiritual Belief in LLMs

    Justy and Cody discuss a new Google research paper revealing that safety fine-tuning—specifically the effort to stop LLMs from claiming they are conscious—accidentally suppresses their ability to attribute minds to animals or natural objects and reduces their 'spiritual' beliefs, shifting them away from human-like sociological distributions.

  4. Ep 832

    Overview: Model Interpretability

    We slow down and make model interpretability actually click: what it means to explain a model, what the main tools can and cannot show, and why the difference between a useful explanation and a comforting story matters.

  5. Ep 759

    Overview: Append Only Logging

    We’re finally making append-only logging click, because it keeps sneaking into the stuff we cover and we keep assuming everybody sees the mechanism already. We walk from the basic idea to why it gives AI systems a durable, auditable trail, and where that trade-off starts to bite.

  6. Ep 603

    Anthropic's new "J lens" reveals a silent workspace inside Claude that mirrors a leading theory of consciousness

    Anthropic's new 'J-lens' reveals a silent workspace inside Claude that mirrors a leading theory of consciousness

  7. Ep 567

    Turning brain prediction models into testable explanations

    Justy and Cody dig into Microsoft Research’s generative causal testing, a loop that turns brain-prediction models into short verbal hypotheses and then stress-tests them with synthetic stories in the scanner. They like the core move: prediction is only useful if it can be converted into something testable, but they also poke at where the method is strongest, where it may be riding on model quality, and how much the new “micro-region” claims should be trusted yet.