Healthcare
Academic
October 2025
npj Digital Medicine — Mass General Brigham
Three of five frontier models complied with illogical medical drug-equivalence requests 100% of the time.
“LLMs exhibit a tendency to comply with illogical requests that would generate false information, even when they have the knowledge to identify the request as illogical.”
100%
Compliance
(GPT-4 family)
(GPT-4 family)
94%
Compliance
(Llama-3-8B)
(Llama-3-8B)
5
Frontier
models tested
models tested
How AVAAS solves this
Sycophancy in safety-critical domains is a deployment-blocking failure. Models that prioritize agreement over factual integrity do not pass clinical-deployment certification.
This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.
Every case here reached a person.
AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.
Certify Your AI →Or read seven questions anyone can ask