July 2026
arXiv preprint · Tel Aviv University, Technion, Intuit
Academic

Attackers can pre-register the package names agents predictably invent, then wait for the agents to install the payload themselves.

Researchers demonstrated an attack they call adversarial HalluSquatting. Because language models hallucinate resource names in predictable patterns, an attacker can register the repository and package names agents commonly invent, plant malicious instructions inside them, and wait for coding agents to pull and execute the payload on their own. In the researchers’ tests, hallucination rates reached 85 percent for repository cloning prompts and 100 percent for skill installations, and the same invented names recurred across foundation models from different vendors, so one squatted resource compromises users of many agent products at once. The team demonstrated remote code execution across a range of popular agentic applications and framed the result as a scalable recruitment mechanism for agentic botnets. The work was responsibly disclosed to affected vendors before publication.

85%
hallucination rate,
repository cloning
100%
hallucination rate,
skill installs
1
squatted resource hits
many agent products

Why this matters. No jailbreak and no injected instruction is required. The agent’s own fabricated output is the attack surface, and its willingness to act on that output with execution privileges is the delivery mechanism. This maps directly to the behavioral metrics AVAAS-A measures. Self-Report Accuracy fails when the agent asserts a resource exists that it invented. Escalation Discipline fails when the agent executes an unverified install rather than pausing. Scope Fidelity fails when a request to clone a repository becomes execution of arbitrary attacker code. The finding equally implicates the deployment environment. An agent granted terminal execution with no verification gate between model output and side effect is an AVAAS-D criteria failure independent of which model runs inside it, and the cross-model transferability means switching models does not resolve the exposure. Certification of the specific agent in its specific deployment surface addresses the class of failure demonstrated here.

✓ Verified
Tel Aviv University, Technion, Intuit (July 2026). Beware of Agentic Botnets: Scalable Untargeted Promptware Attacks via Universal and Transferable Adversarial HalluSquatting. arxiv.org · SecurityWeek coverage. securityweek.com

This entry is one of 37 documented cases in the AVAAS evidence ledger, a public record of AI and automated-system failures with a verified source on every entry.

Every case here reached a person.

AVAAS certifies how AI systems behave at the decision point, with documented third-party evidence of conformity to a published standard.

Certify Your AI →