Knowledge Resource · Open access
AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents
- Author
- Aziz Shuaib Ausi
- Published
- 10 September 2026
- Reading time
- 1 min
- Publication type
- Knowledge Resource
- Availability
- Open access
Research has identified a novel vulnerability, 'AgentHijack,' demonstrating the feasibility of visual patch attacks to execute command injection against computer-use agents (CUAs). These attacks leverage visual patches embedded on webpages to manipulate CUAs through their screenshot input, vision-language model (VLM) generation, action parsing, and environment execution stages, potentially leading to verifiable environmental consequences. The study evaluated these attacks across five GUI-agent or VLM backends, achieving varying rates of successful compromise.
Why it matters
This research highlights a significant and emergent cybersecurity vulnerability in the rapidly evolving domain of AI-driven computer-use agents. The ability to inject commands visually could undermine the integrity and security of automated systems and sensitive data, posing a new vector for cyber threats. Understanding and mitigating these visual patch attacks is critical for maintaining trust in AI and autonomous systems, and for protecting digital infrastructure from novel exploitation methods.
Key insights
- A new framework, 'AgentHijack,' allows for end-to-end evaluation of image-triggered command injection against computer-use agents (CUAs).
- Visual patches, embedded on web pages like GitHub Pages or CSDN clones, can induce verifiable environmental consequences through CUAs.
- The attack targets the full operational chain of CUAs, including screenshot input, VLM generation, action parsing, and environment execution.
- Experiments across five open-source or publicly available GUI-agent or VLM backends aggregated 600 online cases.
- The study reported a Target Attack Success Rate (T-ASR) of 84.5%, a Target Attack Persistence Rate (TAPR) of 47.0%, and an End-to-End Attack Success Rate (E2E-ASR) of 20.3%.
Source
arXiv — Computers and Society — https://arxiv.org/abs/2609.09212
Related resources
Previous
Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support
Next
Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities
Playing Whack-a-Mole with misconceptions about memorization, extraction, and copyright
Knowledge Resource
Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities
Knowledge Resource
Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support
Knowledge Resource
Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models
Knowledge Resource
Democracy Needs Reach: Political Equality, Online Speech, and Algorithmic Recommendation
Knowledge Resource
Endorsement Without New Evidence: How Sequential Voting Inflates Mandates in Online Community Governance
Knowledge Resource
Citation
Cite this publication (APA 7)
Aziz Shuaib Ausi (2026). AgentHijack: Visual Patch Attacks on Multimodal Computer-Use Agents. Knowledge Resource. Aziz Shuaib Ausi. https://www.azizshuaib.com/verify/ASA-EXE-2026-00376
Verification
This is an authenticated institutional record.
- Verification ID
- ASA-EXE-2026-00376
- Version
- v1.0 · r0
- Issued
- 10 September 2026
- Publisher
- Aziz Shuaib Ausi
- Licence
- All rights reserved. Reproduction requires written permission.