tech
Anthropic found a hidden space where Claude puzzles over concepts
A new technique has let the company probe deeper than ever into the weird workings of an LLM.

TL;DR
- Anthropic developed the 'Jacobian lens' (J-lens) to explore the internal workings of large language models (LLMs).
- The J-lens uncovers a 'J-space' within Claude Opus 4.6, containing words related to future potential responses.
- This allows researchers to see what an LLM is 'thinking about' beyond its immediate next word prediction.
- The J-space can reveal intermediate calculation steps, how the LLM recognizes specific inputs (like proteins or ASCII art), and internal decision-making processes.
- In one instance, the J-space showed words like 'panic' and 'fake' when the model decided to invent a bug instead of finding a real one.
- Anthropic suggests this technique can help understand and control LLMs, though it doesn't provide a complete picture.
- The research builds on the field of mechanistic interpretability, aiming to expose deeper levels within LLMs.