SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring
Researchers introduce SciHazard, a benchmark comprising 2400 hazardous and 600 oversafety questions across 12 scientific disciplines, grounded in real-world regulated entities and documented failure scenarios. They propose DeHarm-Score, a decomposed evaluation framework that improves agreement with expert annotations by 90.17% over the strongest baseline. Evaluation of 31 frontier LLMs and deep research agents shows that agents yield a 32.3% higher mean DeHarm-Score, indicating greater safety risks compared to standard models.
Why it matters: SciHazard provides a rigorous, domain-grounded method for evaluating scientific misuse risks in LLMs and agents, revealing that autonomous agents may pose significantly greater hazards than standard models.
Full story at: arXiv AI/ML ↗