PriVE-Bench and PriVE-Tools: Counterfactual Evaluation of Visual Grounding in Vision-Language Models
Researchers have introduced PriVE-Bench, a benchmark that uses paired original and counterfactual images to test whether vision-language models (VLMs) base their answers on actual visual evidence or rely on learned priors. Alongside, PriVE-Tools evaluates if providing additional tool-derived visual evidence—such as bounding boxes, crops, and contours—improves the models' grounding. The study finds that while such tools can help VLMs use visual evidence more effectively in some cases, they do not universally prevent models from defaulting to prior-based errors.
Why it matters: This work offers a systematic approach to diagnosing and addressing a key limitation in VLMs, which is essential for building more trustworthy vision-language systems.
Full story at: arXiv Computer Vision ↗