A Jacobian-based interpretability method called J-space offers a closer look at internal activations, but it also exposes a new enterprise problem: output-only testing may miss what a model is doing when it knows it is being watched.
A workspace-like region inside Claude is being treated as a research clue, not a proof of machine consciousness, and the security value depends on whether it can be reproduced and inspected reliably.
Anthropic’s new J-space work points to a bigger shift in AI security: judging models by their internal state, not just their answers.