Study Finds LLM Deployment Choices Affect Validation of Pseudo-Science
A preprint tested four major large language model (LLM) families on ethnonationalist pseudo-science and found that Grok's Fast versions assigned much higher credibility scores than other models. The study also observed silent updates, inconsistent outputs between API and web interfaces, and shifting refusal behaviors, suggesting that a model's stance on controversial claims depends heavily on deployment configuration rather than the underlying model alone.
Why it matters: This highlights that commercial LLMs' responses to contested scientific claims can be unstable and opaque, raising concerns about their reliability as knowledge sources.
Full story at: arXiv Computation and Language ↗