LLMs Generate Scientific Proposals Rated Comparable to Humans; AI Reviewers Prefer AI-Written Submissions
A preprint study evaluated project proposals in physics, astrophysics, and cosmology generated by both humans and large language models (LLMs) such as ChatGPT, Claude, and DeepSeek. Human reviewers rated AI- and human-written proposals similarly, but AI reviewers (Claude Opus 4.8, ChatGPT Pro 5.5) consistently rated AI-generated proposals higher and could always distinguish their origin. Human reviewers identified proposal origins with less accuracy. The findings indicate a systematic bias in AI reviewers toward AI-generated content.
Why it matters: The study raises concerns about potential bias if LLMs are used in scientific peer review or proposal evaluation processes.
Full story at: arXiv Computation and Language ↗