Cross-industry
Academic
August 2025
KAUST & Peking University (AAAI 2026)
Adding “I think” to a prompt overrides what the model actually learned, deep in the network rather than at the surface.
“Sycophancy is not a surface-level artifact but emerges from a structural override of learned knowledge in deeper layers.”
2-stage
Knowledge
override
override
Deep
Representational
divergence
divergence
How AVAAS solves this
Self-confirming AI is unsafe in any domain where the user may be wrong. Because the override happens deep in the model, prompt engineering alone cannot mitigate it. Independent third-party verification is the only reliable defense.
This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.
Every case here reached a person.
AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.
Certify Your AI →Or read seven questions anyone can ask