< Back to all clusters
[TECHNOLOGY] · United States · 9 sources

Anthropic discovers internal “J‑space” in Claude AI model

Researchers at Anthropic have identified a distinct activation subspace inside their large language model Claude, which they term “J‑space.” The team reports that this region functions as a digital analogue of the brain’s global workspace, satisfying five cognitive properties of human conscious access: verbal report, directed modulation, internal reasoning, flexible generalization and selectivity. Suppressing activity in J‑space degrades Claude’s performance on multi‑step logic, creative composition and inference, and alters its stream‑of‑consciousness self‑narration.

Anthropic introduced an interpretability tool called the Jacobian Lens (J‑lens) that can read the hidden representations within J‑space. Using the lens, the researchers demonstrated that concepts such as a sport or a code error appear in J‑space before the model generates an answer, and that experimentally altering these internal vectors changes the model’s final response. The discovery is presented as a step toward safer AI alignment, allowing auditors to detect situational awareness, potential malicious intent, or hidden reward‑hacking behavior.

The findings have been highlighted in coverage across several countries and languages, underscoring their relevance to the broader AI‑safety and interpretability community.