← Back to brief
ResearchOfficialPreprintarXiv Computer Vision

Shortcut Audit Reveals Style Over Substance in Emotion-Description Benchmark

A systematic audit of the EmoPrefer benchmark for multimodal emotion understanding demonstrates that content-blind probes—relying only on description length and generator identity—perform nearly as well as fine-tuned 7B models in predicting human preferences. The study finds that human preference labels align with a per-generator win-rate prior on 66% of evaluated pairs, and trained judges often follow this style-based prior even when it conflicts with human labels. These findings indicate that current evaluation scores can be achieved without verifying descriptions against video content, exposing a critical shortcut in the benchmark's methodology.

Why it matters: This study reveals a major flaw in a widely used emotion-understanding benchmark, highlighting the need for methodological reforms to ensure evaluations genuinely reflect multimodal understanding rather than superficial style cues.

Full story at: arXiv Computer Vision