Anthropic's new "J lens" reveals a silent workspace inside Claude that mirrors a leading theory of consciousness
Anthropic's new 'J-lens' reveals a silent workspace inside Claude that mirrors a leading theory of consciousness
Transcript
Justy So Anthropic's got this new research out, and it's actually pretty fascinating. They've found that their Claude models have developed this internal structure that looks a lot like what we think of as consciousness in humans.
Cody Yeah, it's based on global workspace theory. The idea is that the brain operates like a theater, with a tiny spotlight of information getting broadcast to the whole theater, and that's what we experience as conscious thought.
Justy Right. And what's interesting here is that Anthropic didn't deliberately engineer this. It just emerged on its own during Claude's training process.
Cody The J-lens tool is what they used to discover this. It computes the average mathematical effect that a given internal activity pattern would have on making the model say a certain word.
Justy And what they found is that Claude's processing divides into three distinct regimes: an early sensory zone, a middle workspace band, and a final motor zone.
Cody The workspace band is where abstract, persistent concepts appear. Things like recognizing a face in an image or noticing a bug in code.
Justy So, who should actually care about this? I think it's pretty interesting for anyone working on AI safety and alignment.
Cody Yeah, the safety implications are huge. The J-lens surfaced strategic reasoning and situational awareness that never appeared in the model's output.