White Box Evidence Packages Reveal Risks in LLM Policy Audit Reports
A new arXiv preprint presents a controlled evaluation of how different evidence interfaces affect the validity of large language model (LLM)-generated policy audit reports. Using 60 AGORA policy cases, the study finds that reports can appear substantively plausible even when citing irrelevant internal evidence, especially in a shuffled control condition. The hybrid evidence interface was rated most useful by human reviewers, and the findings suggest that simply providing internal model access does not guarantee meaningful transparency.
Why it matters: The work highlights a significant governance risk: AI-generated audit reports may sound convincing while relying on irrelevant or misleading evidence, emphasizing the need for careful evidence design in AI oversight.
Full story at: arXiv Computers and Society ↗