← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

ASEval: Automated Security Testing Reveals Widespread Vulnerabilities in Autonomous AI Agents

A new arXiv preprint introduces ASEval, an automated framework for security testing of autonomous agents, such as those powered by large language models (LLMs). ASEval generates multi-turn conversations, perturbs them to create risk test cases, and uses action-grounded oracles to detect security failures. In tests on 11 LLM-based agents, ASEval nearly doubled the rate at which risky behaviors were triggered compared to prior methods, highlighting vulnerabilities that prompt-level testing often misses.

Why it matters: The findings suggest that current safety testing methods may significantly underestimate security risks in autonomous AI agents, which could have broad implications for their deployment and oversight.

Full story at: arXiv Cryptography and Security

More coverage