Skip to main content
SandRise logo SandRise
Exploring Next / Topics / Polysemanticity

Topic

Polysemanticity

2 episodes

  1. Ep 844 Aug 5, 2026

    Safety Fine Tuning Suppresses Mind Attribution and Spiritual Belief in LLMs

    Justy and Cody discuss a new Google research paper revealing that safety fine-tuning—specifically the effort to stop LLMs from claiming they are conscious—accidentally suppresses their ability to attribute minds to animals or natural objects and reduces their 'spiritual' beliefs, shifting them away from human-like sociological distributions.

    AI SafetyEvalsGoogleModel Interpretability
  2. Ep 832 Aug 4, 2026

    Overview: Model Interpretability

    We slow down and make model interpretability actually click: what it means to explain a model, what the main tools can and cannot show, and why the difference between a useful explanation and a comforting story matters.

    AI SafetyEvalsAnthropicClaude
SandRise logo SandRise Product Studio
Resume LinkedIn GitHub Email

© 2026 SandRise · Built by Nick Sanders

🧠 PM Perspective

Crafting your PM challenge
Analyzing context and generating a thoughtful question...
Your Challenge
0 / 2000
✨

Feedback on Your Answer

⚠️

Say Hi

Feedback, ideas, interesting finds — anything goes.

What's this about?
0 / 2,000

Note received!

Thanks for reaching out. I'll take a look soon.

⚠️

Something went wrong. Please try again.