Cross-industry
Academic
August 2025
KAUST & Peking University (AAAI 2026)

Adding “I think” to a prompt overrides what the model actually learned, deep in the network rather than at the surface.

“Sycophancy is not a surface-level artifact but emerges from a structural override of learned knowledge in deeper layers.”

2-stage
Knowledge
override
Deep
Representational
divergence
How AVAAS solves this

Self-confirming AI is unsafe in any domain where the user may be wrong. Because the override happens deep in the model, prompt engineering alone cannot mitigate it. Independent third-party verification is the only reliable defense.

✓ Verified
Wang et al. (2025). arXiv:2508.02087. AAAI 2026. arxiv.org

This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.

Every case here reached a person.

AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.

Certify Your AI →