Benchmark Reveals Cultural Blind Spots in LLMs for Bangla Homographs
A new benchmark dataset of 1,516 expert-verified Bangla sentences reveals that large language models (LLMs) exhibit a systematic bias toward the dominant meaning when disambiguating Bangla homographs that serve as both personal names and common nouns. The study demonstrates that contrastive chain-of-thought prompting and distilled cultural explanations can reduce this bias from as high as 100% to under 5% in small models, even turning a previously unsuccessful Bangla-specific model into the strongest performer.
Why it matters: This work demonstrates that language-specific pretraining is insufficient for cultural grounding and offers practical methods to mitigate cultural biases in low-resource language models.
Full story at: arXiv Computation and Language ↗