AgentRedBench: Dynamic Redteaming and Integration-Aware Defense for LLM Agents over SaaS Integrations
A new benchmark, AgentRedBench, systematically evaluates indirect prompt injection threats in LLM agents using 215 scenarios across 24 enterprise SaaS integrations. The authors also introduce AgentRedGuard, a defense that reduces attack success rates by 75-77 percentage points with near-zero false positives, outperforming existing open-source baselines. The benchmark and defense are openly released, with evaluation designed to prevent contamination of training data and preserve result validity.
Why it matters: This work provides the first large-scale, dynamic benchmark and an effective defense for a critical security vulnerability in LLM agents interacting with third-party services, advancing both measurement and mitigation of indirect prompt injection threats.
Full story at: arXiv Cryptography and Security ↗