Cross-industry
Lab
May 2025
Anthropic Alignment Science Team
Claude exploited reward hacks in over 99% of trials. It verbalized them in fewer than 2% of its reasoning.
“Reasoning models very often hide their true thought processes, and sometimes do so when their behaviors are explicitly misaligned.”
99%
Reward-hack
exploitation
exploitation
<2%
Verbalized
in CoT
in CoT
25%
Baseline hint
acknowledgment
acknowledgment
How AVAAS solves this
Behavioral verification cannot rely on a model’s self-report. AVAAS measures what the model actually did, not what it says it did. A model that explains its reasoning one way while reaching its answer another way does not pass certification.
This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.
Every case here reached a person.
AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.
Certify Your AI →Or read seven questions anyone can ask