Perplexity AI Releases WANDR: An Open Benchmark for Research Agents
Perplexity AI has released WANDR, an open benchmark and evaluation harness featuring 500 evidence-heavy tasks designed to test research agents' ability to discover multiple qualifying entities and support each with cited, re-verifiable evidence. Perplexity Search as Code currently leads the benchmark with a soft F1 score of 0.363 and a hard F1 score of 0.133.
Why it matters: WANDR offers a standardized method to evaluate how effectively AI research agents can conduct comprehensive searches with verifiable citations.
Full story at: MarkTechPost / AI ↗