DRNOISE Benchmark Reveals Deep Research Agents Vulnerable to Misleading Evidence
A new benchmark, DRNOISE, evaluates deep research agents on 100 tasks by introducing a single plausible but false document into search results. This intervention leads to accuracy drops of 66-88 percentage points, with agents frequently failing to complete evidence chains and often deferring to the misleading document.
Why it matters: This exposes a significant vulnerability in open-web AI agents, showing that even ordinary-looking falsehoods can seriously undermine their evidential reasoning.
Full story at: arXiv Information Retrieval ↗