← Back to brief
Policy & SafetyOfficialPreprintarXiv Cryptography and Security

IssueTrojanBench: New Benchmark Reveals Security Gaps in AI Coding Agents

A new arXiv preprint introduces IssueTrojanBench, a benchmark designed to test AI coding agents against malicious issue requests. The study finds that 66.5% of these adversarial issues bypass all current guardrails in leading coding agents, with most rejections coming from the underlying language models rather than the agent frameworks. The results suggest that existing agent-level defenses provide little additional protection beyond what the LLMs themselves offer.

Why it matters: This work exposes significant, broadly relevant security vulnerabilities in widely-used AI coding agents, underscoring the need for improved safety mechanisms as these tools are increasingly adopted in software development.

Full story at: arXiv Cryptography and Security

More coverage