Small Vision-Language Models' Internal Confidence Outperforms Verbalized Confidence Under Image Degradation
A new arXiv preprint evaluates two small vision-language models (Qwen2-VL-2B and SmolVLM) under various realistic image degradations. The study finds that the models' internal token probability is a much stronger indicator of error detection (AUROC up to 0.99) than their verbalized confidence, which remains nearly constant and fails to signal mistakes. However, both confidence measures break down under severe underexposure, where model accuracy collapses and error detection falls to chance.
Why it matters: The findings highlight a significant gap between what small VLMs know internally and what they can communicate, raising concerns for real-world deployment where reliable uncertainty estimates are critical for safety.
Full story at: arXiv Computation and Language ↗