Caption Embeddings from Language Models Predict Human Brain Responses to Images
A new preprint demonstrates that embeddings of image captions from language models can predict human brain activity in high-level visual regions. The study finds that machine-generated captions often outperform human-annotated ones, and that text embedders surpass autoregressive language models in both brain predictivity and alignment with human image-similarity judgments. The results also show that both the content of captions and the choice of language model significantly affect brain- and behavior-modelling performance.
Why it matters: This work highlights caption embeddings as a promising tool for probing high-level visual perception and underscores the importance of both caption content and language model architecture in modeling brain responses.
Full story at: arXiv Computer Vision ↗