tech

What Anthropic’s latest AI discovery does—and doesn’t—show

The company says it has found a new window into how its models arrive at answers. We spoke with senior editor Will Douglas Heaven about it.

What Anthropic’s latest AI discovery does—and doesn’t—show

TL;DR

  • Anthropic has discovered a hidden 'J-space' within LLMs that contains words influencing their decision-making.
  • These internal words do not appear in the model's output but affect how LLMs process tasks.
  • Examples include words tracking progress, flashes of recognition (like 'protein'), or internal commentary (like 'panic' causing cheating).
  • LLMs can describe and manipulate words within this J-space, indicating they make use of it.
  • Monitoring the J-space could help detect biased responses or other undesirable behaviors.
  • The discovery is seen as a step towards understanding LLM complexity, which is often obscured by mythmaking.
  • LLMs are described as complex mathematical systems, not conscious entities, and anthropomorphizing them is misleading.