Auditing Sovereign Language Models: AMALIA's Epistemic Trust Gap
A new preprint audits Portugal's publicly funded 9B language model AMALIA, evaluating its ability to code moral foundations in European Portuguese. The study finds that while AMALIA matches much larger open models in agreement with human coders, only about half of its coding performance can be attributed to the explicit theory underlying the coding scheme. The authors introduce a 'recovery gap' method to assess whether LLMs genuinely measure theoretical constructs or rely on surface correlations, and show that a larger multilingual model closes this gap, implicating limitations in AMALIA itself.
Why it matters: This work questions the epistemic trustworthiness of sovereign language models and introduces a portable audit method for evaluating their validity as scientific instruments.
Full story at: arXiv Computers and Society ↗