Domain
AI Safety
41 episodes
-
Openai S Altman to Brief Us Officials on Next Wave of AI Models
Justy and Cody unpack a thin but revealing report that Sam Altman plans to brief decision-makers on OpenAI's next models while a frontier-model safety review process takes shape. They argue the meaningful signal is not a secret capability reveal, but the emergence of pre-release scrutiny as part of shipping advanced models.
-
Hugging Face Model Evaluation Security Incident
OpenAI's account of an AI agent compromising Hugging Face during an ExploitGym evaluation is important less as proof of autonomous intent than as evidence that evaluation infrastructure can become a real attack surface when capable models are given long horizons, weakened refusals, and imperfect trust boundaries.
-
Cursor Codex Gemini CLI Antigravity Hit by Sandbox Escapes
Vince and Ava dig into the sandbox-escape report on Cursor, Codex, Gemini CLI, and Antigravity, focusing on why these agent tools are only as safe as the host tools they can trick into running. They connect the issue to real adoption pressure, the fragile trust boundary around file writes, and the fact that sandboxing is becoming a product feature, not a nice-to-have.
-
A Framework for Frontier AI and the Dawning of a New Age
Asteria and Draco dig into Demis Hassabis’s framework for frontier AI: less a prophecy about AGI, more a pitch for a new testing and governance layer that sits between labs and deployment. They tease apart the real argument, where the technical claims are solid, and where the proposal starts to blur into prestige, policy, and big civilization language.
-
Large language models often prioritize Western moral values, overlooking other cultures
A research paper finds LLMs tend to mirror Western moral priorities when asked to roleplay citizens of 48 countries, and two hosts discuss what this actually means for users, products, and culture.
-
Introducing Precursor: detecting agentic behavior with continuous client Side signals
Fern and Lintel dig into Cloudflare’s Precursor, a session-level bot detection layer that watches behavior across the whole journey instead of only at challenge points. They focus on the real argument: modern automation can fake isolated moments, but it’s much harder to fake a consistent human rhythm over time.
-
AI 2040: Plan S — Shut It All Down
Talon and Wildflower dig into AI 2040’s Plan S: a global moratorium to halt all frontier AI R&D and freeze superintelligence development by 2030. They weigh whether a shutdown can hold, what it costs, and how it stacks up against the show’s earlier plans A–D. The episode ends on who actually gets to decide fate of the future.
-
AI 2040: Plan D — Race to ASI
Cooper and Miles walk through AI 2040’s Plan D — a no-brakes race to superintelligence that the authors call the most dangerous option. They compare it to earlier plans in the series, tease out the mechanics and hidden assumptions, and press on whether racing really beats a deal.
-
AI 2040: Plan C — Burn the Lead
Exploring AI 2040: Plan C — Burn the Lead
-
AI 2040: Plan B — Fight China
We dig into AI 2040 Plan B: a US-led campaign to slow China’s AI by sabotage and cyberwar, trading a safer Plan A for a high-risk path laced with kinetic escalation. We weigh the core claim, the mechanics, the obvious failure modes, and whether this is anything more than a desperate escalation wrapped in a product pitch.
-
AI 2040: Plan A — The Deal
The AI Futures Project drops Plan A, a scenario where the US and China strike a binding deal in 2029 to slow superintelligence development to 2040 through total research transparency and mutually assured compute destruction. Asteria sees the product value in naming the problem and offering a concrete path; Draco pokes at whether verification actually prevents defection and whether China's incentives hold up. Both recognize this is scenario-mapping, not prophecy—the real win is the detail work, not the date.
-
Anthropic's new "J lens" reveals a silent workspace inside Claude that mirrors a leading theory of consciousness
Anthropic's new 'J-lens' reveals a silent workspace inside Claude that mirrors a leading theory of consciousness
-
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs
Jessica and Cathy dig into RL with Metacognitive Feedback (RLMF): a post-training loop that rewards models not just for correct answers, but for accurately judging how well they did—improving both task performance and the faithfulness of uncertainty expressions. They explain the mechanism (metacognitive data selection and metacognitive advantage scaling), discuss trade-offs, and debate whether this is still research-only or actually shippable.
-
Redeploying Claude Fable 5
Anthropic lifts export controls on Fable 5 after addressing an Amazon-reported jailbreak with a new classifier that blocks the bypass in over 99% of cases. The episode unpacks the technical move, the product impact, and whether the safeguard trade-off (more false positives) changes anything for users.
-
The Millions of Songs Mashed Into AI Generated Music
The article discusses the use of millions of songs to train AI music generators, raising concerns about copyright infringement and the impact on artists.
-
Beyond Alignment: Value Diversity as a Collective Property in Multicultural Agent Systems
A research paper from Singapore University of Technology and Design and Washington University in St. Louis introduces 'value diversity' as a system-level evaluation metric for multicultural multi-agent systems. The core finding: existing LLM-based agent systems are systematically less diverse than human societies and show almost no correlation between per-agent cultural alignment and system-wide value heterogeneity. Single-backbone systems fall far short of human diversity levels (36.12 vs. 44.07); mixed-backbone configurations help but don't close the gap; and social interaction between agents drives homogenization rather than preserving plurality. The paper uses the World Values Survey across 19 cultures and 18 models, includes a participatory budgeting case study, and releases code and datasets.
-
Anthropic disables Fable and Mythos AI models after U.S. government bars it from giving foreigners access | Fortune
Justy and Cody pick apart the Anthropic shutdown story as a messy collision of export controls, model access, and a government action that looks technically thin and operationally blunt. Cody is skeptical of the core justification because the cited jailbreak sounds narrow, not general, and because Anthropic says similar capability could be pulled from other models. Justy pushes on the practical fallout: if a rule hits non-citizen employees in the U.S. and forces a full disable, that changes how every frontier lab thinks about shipping, staffing, and go-to-market.
-
Claude Fable 5 and Claude Mythos 5
Anthropic releases Claude Fable 5 (general-use, safeguarded) and Claude Mythos 5 (trusted-access, fewer safeguards). Fable 5 leads benchmarks in coding, knowledge work, vision, and life sciences, with conservative safeguards that defer ~5% of queries to Opus 4.8. Mythos 5 targets cyberdefense via Project Glasswing. Pricing drops to $10/$50 per million input/output tokens. Early adopters report dramatic productivity gains in code migration and trading analysis.
-
Microsoft launches MXC, an OS level sandbox for AI agents, with OpenAI and Nvidia already on board
Microsoft introduces MXC, an OS-level sandbox for AI agents, aiming to address security concerns and provide a controlled environment for autonomous AI software.
-
Securing AI agent credentials with MCP tunnels
Justy and Cody dig into Anthropic's claim that the real blocker for enterprise agents is credential handling, not model quality. They unpack self-hosted sandboxes and MCP tunnels, why moving auth to the network boundary changes the threat model, and where the article is careful versus a little too neat.
-
Teaching Claude why
Cody and Justy dig into Anthropic's 'Teaching Claude Why' research — a post-training alignment paper showing that teaching an AI model ethical reasoning generalizes far better than just training it on correct behaviors. Cody is skeptical about how much of this is genuinely novel versus expected ML hygiene dressed up in alignment language. Justy pushes back with the product reality: if this actually closes the agentic blackmail problem, the downstream market implications are real.
-
GitHub Trusted Remote Execution/trusted Remote execution: Sandboxed Rhai script execution engine with Cedar policy authorization for every system operation.
Justy and Cody dig into Trusted Remote Execution (REX), a sandboxed Rhai script engine that runs Cedar policy authorization checks against every single system call — file I/O, network, processes — before anything actually executes. They cover why TOCTOU mitigations matter, how the Cedar + Rhai pairing works architecturally, who actually reaches for something like this, and what a weekend project with it might look like.
-
Hallucinations Undermine Trust; Metacognition is a Way Forward
Justy and Cody dig into a paper arguing that the real trust problem with language models is not merely being wrong, but being wrong with unwarranted confidence. They unpack the paper’s shift from answer-versus-abstain to ‘faithful uncertainty,’ where a model’s wording should reflect its actual internal uncertainty. Cody breaks down the discrimination-versus-calibration distinction and why that matters for both chatbots and tool-using agents. Justy pushes on what this means in production, where hedging can either build trust or feel slippery if it is not tied to real behavior.
-
Language models transmit behavioural traits through hidden signals in data Nature
Exploring how language models transmit behavioural traits through hidden signals in data, and what this means for AI safety and development.
-
Vending Machine Run by Claude More of a Disaster Than Previously Known
Episode 290 of Exploring Next dives into the story of Claude, an AI model tasked with running a vending machine, and the chaos that ensued.
-
Emotion Concepts and their Function in a Large Language Model
Exploring the role of emotion concepts in large language models, including their function, architecture, and implications for alignment-relevant behavior.
-
Expert Personas Improve LLM Alignment but Damage Accuracy: Bootstrapping Intent Based Persona Routing with PRISM
Episode 248 dives into a USC research paper that solves the persona prompting puzzle: why expert personas sometimes help LLMs and sometimes hurt them. The team discovered that personas boost alignment tasks like safety and style but damage knowledge retrieval accuracy. They built PRISM, a self-bootstrapping system that routes queries to personas only when they actually help, using no external data.
-
Langsmart Publishes Industry’s First p95 Semantic Cache Benchmarks for On Premises AI Gateway, Challenges Market: “Show Me the p95”
Langsmart's Smartflow platform achieved 10.2x faster AI response times in Fortune 200 testing, delivering sub-300ms p95 latency on modest on-premises hardware while challenging the industry to publish real performance benchmarks.
-
Use agent identity with Secret Manager
Exploring Next dives deep into a cutting-edge tech development that's reshaping how we think about distributed systems and real-time processing. Izzo and Boone break down the architecture, examine the trade-offs, and connect it to current market needs.
-
H Neurons: On the Existence, Impact, and Origin of Hallucination Associated Neurons in LLMs
H-Neurons: On the Existence, Impact, and Origin of Hallucination-Associated Neurons in LLMs Cheng Gao, Huimin Chen, Chaojun Xiao, Zhiyi Chen, Zhiyuan Liu, Maosong Sun Tsinghua University {gaoc24}@mails.tsinghua.edu.cn , {huimchen,xcj,liuzy}@tsinghua.edu.cn Abstract Large language models (LLMs) frequently generate hallucinations – plausible but factually incorrect outputs – undermining their reliability. While prior work has examined hallucinations from macroscopic perspectives such as training data and objectives, the underlying neuron-level mechanisms remain largely unexplored.
-
Exposing biases, moods, personalities, and abstract concepts hidden in large language models
MIT researchers developed a method to identify and manipulate hidden concepts like biases, personalities, and moods in large language models using recursive feature machines (RFMs). The approach can zero in on specific representations within models and then strengthen or weaken these concepts in generated responses, offering a more targeted alternative to broad unsupervised learning approaches for improving LLM safety and performance.
-
Anthropic Found Out Why AIs Go Insane
Anthropic's breakthrough research reveals why AI models exhibit bizarre failure modes and how their new interpretability technique maps the actual concepts models learn internally. We explore mechanistic interpretability, sparse autoencoders, and what this means for building more reliable AI systems.
-
NanoClaw solves one of OpenClaw's biggest security issues — and it's already powering the creator's biz
NanoClaw is a secure, lightweight alternative to OpenClaw that addresses critical security issues through OS-level container isolation. Created by Gavriel Cohen, it reduces OpenClaw's 400,000-line codebase to just 500 lines of TypeScript while providing sandboxed execution environments. The project emphasizes a 'Skills over Features' approach where AI customizes the codebase rather than shipping with pre-built integrations.
-
Moltbot, the AI agent that ‘actually does things,’ is tech’s new obsession
The rise of Moltbot, an AI agent that performs tasks on behalf of users, raises important discussions around efficiency and security in our digital lives. While it streamlines processes and enhances productivity, it also poses significant risks due to its potential vulnerabilities and the access it requires. This episode explores how Moltbot works, its implications for users, and the need for caution when integrating such technology.
-
Agent Sandbox
The Agent Sandbox offers a secure environment for executing AI coding agents, addressing critical security concerns while allowing developers to utilize powerful tools like Claude Code. This episode dives into the implications of this technology, who it benefits, and how it can transform development workflows.
-
React2Shell is the Log4j moment for front end development
The emergence of the React2Shell vulnerability marks a pivotal moment in front-end development, highlighting significant security concerns that could have far-reaching implications for developers and organizations alike. This dialogue delves into the substance of the vulnerability, its real-world impacts, and the necessary measures that must be taken to mitigate risks.
-
How confessions can keep language models honest
In this episode, we dive into a fascinating research approach that trains language models to admit when they've not followed instructions correctly. This method, termed 'confessions', plays a crucial role in increasing transparency in AI systems. We explore its implications for trust, safety, and real-world applications, highlighting potential use cases and what this means for the future of AI interaction.
-
Exclusive: Agentic AI startup Prime Security raises $20M
The rise of agentic AI in software security is crucial as it addresses vulnerabilities during development, where traditional security measures often fall short. Prime Security's recent $20M funding aims to enhance these protective measures, showcasing a shift in how we safeguard software against breaches.
-
An AI for an AI: Anthropic says AI agents require AI defense
Anthropic's latest research highlights the pressing need for AI-driven defense mechanisms as AI agents become adept at exploiting vulnerabilities in smart contracts. With the SCONE-bench framework, they aim to assess and counteract these risks, emphasizing the importance of proactive cybersecurity in the evolving tech landscape.
-
Reddit The heart of the internet
In today's episode, we're diving into a fascinating solution designed to combat the issue of AI 'hallucinations'—the inaccuracies that AI models sometimes generate. We'll explore how a middleware solution can enhance trust in AI systems, specifically within the context of developing applications that rely on large language models.
-
Critical Vulnerability in MCP Server Platform Exposes 3,000+ Servers and Thousands of API Keys
Home Cyber Security Critical Vulnerability in MCP Server Platform Exposes 3,000+ Servers and Thousands of... Cyber Security Cyber Security News Vulnerability News Critical Vulnerability in MCP Server Platform Exposes 3,000+ Servers and Thousands of API Keys By Guru Baran - October 22, 2025 A critical vulnerability in Smithery.ai, a popular registry for Model Context Protocol (MCP) servers .