LLMs Show Major Weaknesses in Detecting Their Own Outputs on Short-Answer Educational Tasks
A new arXiv preprint finds that large language models (LLMs) are unreliable at distinguishing their own generated content from human-written responses in short-answer educational tasks. While LLMs perform well at detecting AI-generated code and longer reflective writing, they often misclassify their own short answers as more human-like than authentic student work. The study also shows that prompt variations can significantly affect detection accuracy in reflective writing, but have less impact on programming tasks.
Why it matters: This highlights a critical limitation for using LLMs as automated detectors of AI-generated student work, especially for short-answer formats that are common in education.
Full story at: arXiv Computation and Language ↗