Safety That Does Not Transfer: Cross-Lingual Clinical Correctness Drift in Deployable Medical Language Models
A preprint study evaluated the clinical safety of large language models (LLMs) in both English and Hausa, focusing on locally deployable models (4-9B parameters) and a frontier model. The results show that deployable models experienced a sharp drop in clinical correctness when answering in Hausa (mean score fell from 1.57 to -0.03), while the frontier model maintained high performance (2.00 to 1.75). The study attributes this deficit to the deployable model tier, rather than the language or clinical task.
Why it matters: This highlights a critical safety gap for low-resource healthcare settings that rely on small, locally-run language models, as clinical correctness may not transfer across languages.
Full story at: arXiv Computation and Language ↗