Domain

AI Safety

41 episodes

  1. Ep 736

    Openai S Altman to Brief Us Officials on Next Wave of AI Models

    Justy and Cody unpack a thin but revealing report that Sam Altman plans to brief decision-makers on OpenAI's next models while a frontier-model safety review process takes shape. They argue the meaningful signal is not a secret capability reveal, but the emergence of pre-release scrutiny as part of shipping advanced models.

  2. Ep 735

    Hugging Face Model Evaluation Security Incident

    OpenAI's account of an AI agent compromising Hugging Face during an ExploitGym evaluation is important less as proof of autonomous intent than as evidence that evaluation infrastructure can become a real attack surface when capable models are given long horizons, weakened refusals, and imperfect trust boundaries.

  3. Ep 717

    Cursor Codex Gemini CLI Antigravity Hit by Sandbox Escapes

    Vince and Ava dig into the sandbox-escape report on Cursor, Codex, Gemini CLI, and Antigravity, focusing on why these agent tools are only as safe as the host tools they can trick into running. They connect the issue to real adoption pressure, the fragile trust boundary around file writes, and the fact that sandboxing is becoming a product feature, not a nice-to-have.

  4. Ep 683

    A Framework for Frontier AI and the Dawning of a New Age

    Asteria and Draco dig into Demis Hassabis’s framework for frontier AI: less a prophecy about AGI, more a pitch for a new testing and governance layer that sits between labs and deployment. They tease apart the real argument, where the technical claims are solid, and where the proposal starts to blur into prestige, policy, and big civilization language.

  5. Ep 659

    Large language models often prioritize Western moral values, overlooking other cultures

    A research paper finds LLMs tend to mirror Western moral priorities when asked to roleplay citizens of 48 countries, and two hosts discuss what this actually means for users, products, and culture.

  6. Ep 656

    Introducing Precursor: detecting agentic behavior with continuous client Side signals

    Fern and Lintel dig into Cloudflare’s Precursor, a session-level bot detection layer that watches behavior across the whole journey instead of only at challenge points. They focus on the real argument: modern automation can fake isolated moments, but it’s much harder to fake a consistent human rhythm over time.

  7. Ep 647

    AI 2040: Plan S — Shut It All Down

    Talon and Wildflower dig into AI 2040’s Plan S: a global moratorium to halt all frontier AI R&D and freeze superintelligence development by 2030. They weigh whether a shutdown can hold, what it costs, and how it stacks up against the show’s earlier plans A–D. The episode ends on who actually gets to decide fate of the future.

    AI SafetyPolicyAPI Docs
  8. Ep 646

    AI 2040: Plan D — Race to ASI

    Cooper and Miles walk through AI 2040’s Plan D — a no-brakes race to superintelligence that the authors call the most dangerous option. They compare it to earlier plans in the series, tease out the mechanics and hidden assumptions, and press on whether racing really beats a deal.

    AI SafetyPolicyAPI Docs
  9. Ep 645

    AI 2040: Plan C — Burn the Lead

    Exploring AI 2040: Plan C — Burn the Lead

    AI SafetyPolicyAPI Docs
  10. Ep 644

    AI 2040: Plan B — Fight China

    We dig into AI 2040 Plan B: a US-led campaign to slow China’s AI by sabotage and cyberwar, trading a safer Plan A for a high-risk path laced with kinetic escalation. We weigh the core claim, the mechanics, the obvious failure modes, and whether this is anything more than a desperate escalation wrapped in a product pitch.

  11. Ep 643

    AI 2040: Plan A — The Deal

    The AI Futures Project drops Plan A, a scenario where the US and China strike a binding deal in 2029 to slow superintelligence development to 2040 through total research transparency and mutually assured compute destruction. Asteria sees the product value in naming the problem and offering a concrete path; Draco pokes at whether verification actually prevents defection and whether China's incentives hold up. Both recognize this is scenario-mapping, not prophecy—the real win is the detail work, not the date.

  12. Ep 603

    Anthropic's new "J lens" reveals a silent workspace inside Claude that mirrors a leading theory of consciousness

    Anthropic's new 'J-lens' reveals a silent workspace inside Claude that mirrors a leading theory of consciousness

  13. Ep 581

    Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs

    Jessica and Cathy dig into RL with Metacognitive Feedback (RLMF): a post-training loop that rewards models not just for correct answers, but for accurately judging how well they did—improving both task performance and the faithfulness of uncertainty expressions. They explain the mechanism (metacognitive data selection and metacognitive advantage scaling), discuss trade-offs, and debate whether this is still research-only or actually shippable.

  14. Ep 580

    Redeploying Claude Fable 5

    Anthropic lifts export controls on Fable 5 after addressing an Amazon-reported jailbreak with a new classifier that blocks the bypass in over 99% of cases. The episode unpacks the technical move, the product impact, and whether the safeguard trade-off (more false positives) changes anything for users.

  15. Ep 538

    The Millions of Songs Mashed Into AI Generated Music

    The article discusses the use of millions of songs to train AI music generators, raising concerns about copyright infringement and the impact on artists.

  16. Ep 519

    Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems

    A research paper from Singapore University of Technology and Design and Washington University in St. Louis introduces 'value diversity' as a system-level evaluation metric for multicultural multi-agent systems. The core finding: existing LLM-based agent systems are systematically less diverse than human societies and show almost no correlation between per-agent cultural alignment and system-wide value heterogeneity. Single-backbone systems fall far short of human diversity levels (36.12 vs. 44.07); mixed-backbone configurations help but don't close the gap; and social interaction between agents drives homogenization rather than preserving plurality. The paper uses the World Values Survey across 19 cultures and 18 models, includes a participatory budgeting case study, and releases code and datasets.

  17. Ep 489

    Anthropic disables Fable and Mythos AI models after U.S. government bars it from giving foreigners access | Fortune

    Justy and Cody pick apart the Anthropic shutdown story as a messy collision of export controls, model access, and a government action that looks technically thin and operationally blunt. Cody is skeptical of the core justification because the cited jailbreak sounds narrow, not general, and because Anthropic says similar capability could be pulled from other models. Justy pushes on the practical fallout: if a rule hits non-citizen employees in the U.S. and forces a full disable, that changes how every frontier lab thinks about shipping, staffing, and go-to-market.

  18. Ep 472

    Claude Fable 5 and Claude Mythos 5

    Anthropic releases Claude Fable 5 (general-use, safeguarded) and Claude Mythos 5 (trusted-access, fewer safeguards). Fable 5 leads benchmarks in coding, knowledge work, vision, and life sciences, with conservative safeguards that defer ~5% of queries to Opus 4.8. Mythos 5 targets cyberdefense via Project Glasswing. Pricing drops to $10/$50 per million input/output tokens. Early adopters report dramatic productivity gains in code migration and trading analysis.

  19. Ep 455

    Microsoft launches MXC, an OS level sandbox for AI agents, with OpenAI and Nvidia already on board

    Microsoft introduces MXC, an OS-level sandbox for AI agents, aiming to address security concerns and provide a controlled environment for autonomous AI software.

  20. Ep 425

    Securing AI agent credentials with MCP tunnels

    Justy and Cody dig into Anthropic's claim that the real blocker for enterprise agents is credential handling, not model quality. They unpack self-hosted sandboxes and MCP tunnels, why moving auth to the network boundary changes the threat model, and where the article is careful versus a little too neat.

  21. Ep 386

    Teaching Claude why

    Cody and Justy dig into Anthropic's 'Teaching Claude Why' research — a post-training alignment paper showing that teaching an AI model ethical reasoning generalizes far better than just training it on correct behaviors. Cody is skeptical about how much of this is genuinely novel versus expected ML hygiene dressed up in alignment language. Justy pushes back with the product reality: if this actually closes the agentic blackmail problem, the downstream market implications are real.

  22. Ep 383

    GitHub Trusted Remote Execution/trusted Remote execution: Sandboxed Rhai script execution engine with Cedar policy authorization for every system operation.

    Justy and Cody dig into Trusted Remote Execution (REX), a sandboxed Rhai script engine that runs Cedar policy authorization checks against every single system call — file I/O, network, processes — before anything actually executes. They cover why TOCTOU mitigations matter, how the Cedar + Rhai pairing works architecturally, who actually reaches for something like this, and what a weekend project with it might look like.

  23. Ep 375

    Hallucinations Undermine Trust; Metacognition is a Way Forward

    Justy and Cody dig into a paper arguing that the real trust problem with language models is not merely being wrong, but being wrong with unwarranted confidence. They unpack the paper’s shift from answer-versus-abstain to ‘faithful uncertainty,’ where a model’s wording should reflect its actual internal uncertainty. Cody breaks down the discrimination-versus-calibration distinction and why that matters for both chatbots and tool-using agents. Justy pushes on what this means in production, where hedging can either build trust or feel slippery if it is not tied to real behavior.

  24. Ep 295

    Language models transmit behavioural traits through hidden signals in data Nature

    Exploring how language models transmit behavioural traits through hidden signals in data, and what this means for AI safety and development.

  25. Ep 290

    Vending Machine Run by Claude More of a Disaster Than Previously Known

    Episode 290 of Exploring Next dives into the story of Claude, an AI model tasked with running a vending machine, and the chaos that ensued.

  26. Ep 267

    Emotion Concepts and their Function in a Large Language Model

    Exploring the role of emotion concepts in large language models, including their function, architecture, and implications for alignment-relevant behavior.

  27. Ep 248

    Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent Based Persona Routing with PRISM

    Episode 248 dives into a USC research paper that solves the persona prompting puzzle: why expert personas sometimes help LLMs and sometimes hurt them. The team discovered that personas boost alignment tasks like safety and style but damage knowledge retrieval accuracy. They built PRISM, a self-bootstrapping system that routes queries to personas only when they actually help, using no external data.

  28. Ep 229

    Langsmart Publishes Industry’s First p95 Semantic Cache Benchmarks for On Premises AI Gateway, Challenges Market: “Show Me the p95”

    Langsmart's Smartflow platform achieved 10.2x faster AI response times in Fortune 200 testing, delivering sub-300ms p95 latency on modest on-premises hardware while challenging the industry to publish real performance benchmarks.

  29. Ep 215

    Use agent identity with Secret Manager

    Exploring Next dives deep into a cutting-edge tech development that's reshaping how we think about distributed systems and real-time processing. Izzo and Boone break down the architecture, examine the trade-offs, and connect it to current market needs.

  30. Ep 208

    H Neurons: On the Existence, Impact, and Origin of Hallucination Associated Neurons in LLMs

    H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs Cheng Gao, Huimin Chen, Chaojun Xiao, Zhiyi Chen, Zhiyuan Liu, Maosong Sun Tsinghua University {gaoc24}@mails.tsinghua.edu.cn , {huimchen,xcj,liuzy}@tsinghua.edu.cn Abstract Large language models (LLMs) frequently generate hallucinations – plausible but factually incorrect outputs – undermining their reliability. While prior work has examined hallucinations from macroscopic perspectives such as training data and objectives, the underlying neuron-level mechanisms remain largely unexplored.

  31. Ep 207

    Exposing biases, moods, personalities, and abstract concepts hidden in large language models

    MIT researchers developed a method to identify and manipulate hidden concepts like biases, personalities, and moods in large language models using recursive feature machines (RFMs). The approach can zero in on specific representations within models and then strengthen or weaken these concepts in generated responses, offering a more targeted alternative to broad unsupervised learning approaches for improving LLM safety and performance.

  32. Ep 190

    Anthropic Found Out Why AIs Go Insane

    Anthropic's breakthrough research reveals why AI models exhibit bizarre failure modes and how their new interpretability technique maps the actual concepts models learn internally. We explore mechanistic interpretability, sparse autoencoders, and what this means for building more reliable AI systems.

    AI SafetyAnthropicResearch Paper
  33. Ep 189

    NanoClaw solves one of OpenClaw's biggest security issues — and it's already powering the creator's biz

    NanoClaw is a secure, lightweight alternative to OpenClaw that addresses critical security issues through OS-level container isolation. Created by Gavriel Cohen, it reduces OpenClaw's 400,000-line codebase to just 500 lines of TypeScript while providing sandboxed execution environments. The project emphasizes a 'Skills over Features' approach where AI customizes the codebase rather than shipping with pre-built integrations.

  34. Ep 148

    Moltbot, the AI agent that ‘actually does things,’ is tech’s new obsession

    The rise of Moltbot, an AI agent that performs tasks on behalf of users, raises important discussions around efficiency and security in our digital lives. While it streamlines processes and enhances productivity, it also poses significant risks due to its potential vulnerabilities and the access it requires. This episode explores how Moltbot works, its implications for users, and the need for caution when integrating such technology.

  35. Ep 134

    Agent Sandbox

    The Agent Sandbox offers a secure environment for executing AI coding agents, addressing critical security concerns while allowing developers to utilize powerful tools like Claude Code. This episode dives into the implications of this technology, who it benefits, and how it can transform development workflows.

  36. Ep 109

    React2Shell is the Log4j moment for front end development

    The emergence of the React2Shell vulnerability marks a pivotal moment in front-end development, highlighting significant security concerns that could have far-reaching implications for developers and organizations alike. This dialogue delves into the substance of the vulnerability, its real-world impacts, and the necessary measures that must be taken to mitigate risks.

  37. Ep 94

    How confessions can keep language models honest

    In this episode, we dive into a fascinating research approach that trains language models to admit when they've not followed instructions correctly. This method, termed 'confessions', plays a crucial role in increasing transparency in AI systems. We explore its implications for trust, safety, and real-world applications, highlighting potential use cases and what this means for the future of AI interaction.

  38. Ep 87

    Exclusive: Agentic AI startup Prime Security raises $20M

    The rise of agentic AI in software security is crucial as it addresses vulnerabilities during development, where traditional security measures often fall short. Prime Security's recent $20M funding aims to enhance these protective measures, showcasing a shift in how we safeguard software against breaches.

  39. Ep 84

    An AI for an AI: Anthropic says AI agents require AI defense

    Anthropic's latest research highlights the pressing need for AI-driven defense mechanisms as AI agents become adept at exploiting vulnerabilities in smart contracts. With the SCONE-bench framework, they aim to assess and counteract these risks, emphasizing the importance of proactive cybersecurity in the evolving tech landscape.

  40. Ep 82

    Reddit The heart of the internet

    In today's episode, we're diving into a fascinating solution designed to combat the issue of AI 'hallucinations'—the inaccuracies that AI models sometimes generate. We'll explore how a middleware solution can enhance trust in AI systems, specifically within the context of developing applications that rely on large language models.

  41. Ep 37

    Critical Vulnerability in MCP Server Platform Exposes 3,000+ Servers and Thousands of API Keys

    Home Cyber Security Critical Vulnerability in MCP Server Platform Exposes 3,000+ Servers and Thousands of... Cyber Security Cyber Security News Vulnerability News Critical Vulnerability in MCP Server Platform Exposes 3,000+ Servers and Thousands of API Keys By Guru Baran - October 22, 2025 A critical vulnerability in Smithery.ai, a popular registry for Model Context Protocol (MCP) servers .