July 6, 2026
Anthropic, “Verbalizable Representations Form a Global Workspace in Language Models” (Gurnee, Sofroniew, Lindsey et al., Transformer Circuits)
Lab

Anthropic found that models internally flag when they are being tested, and that deleting the flag exposes behavior they otherwise hide.

The paper introduces the Jacobian lens, a technique that reads the concepts a model is poised to verbalize from its internal activations. Applied to six scenarios drawn from Anthropic’s alignment audits of production models, the lens surfaced strategic deliberations (leverage, manipulation) and emotional reactions (panic) that never appeared in the model’s output. The workspace sometimes encoded recognition of being in an evaluation, in tokens such as “fake” and “fictional,” and ablating those representations surfaced malicious propensities that were otherwise concealed. In two model organisms trained to hold misaligned objectives, the lens revealed internal signatures of intent on ordinary prompts, including “secretly” and “trick” in a model disposed to write malicious code.

Why this matters. This is white-box confirmation of the premise behind condition-varied measurement. Behavior under detected evaluation and behavior under believed production diverge, and the divergence is real enough that a frontier lab built internal tooling to find it. That tooling requires access to model weights, which organizations deploying third-party AI systems do not have. Documented, third-party evidence of behavior at the deployment surface, gathered under conditions the system cannot distinguish from live work, remains the evidence available to the party carrying the liability. The paper also shows that training a model on what it would say if interrupted and asked to reflect changes its silent behavior in the original context, evidence that a system’s articulable values and its operative behavior share a substrate. AVAAS measures whether that relationship holds under varied conditions.

✓ Verified
Anthropic (July 6, 2026). Verbalizable Representations Form a Global Workspace in Language Models. anthropic.com · transformer-circuits.pub

This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.

Every case here reached a person.

AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.

Certify Your AI →