VCG-Bench: A Unified Benchmark for Structured Diagram Generation and Editing with Vision-Language Models
Researchers introduce VCG-Bench, a benchmark designed to evaluate vision-language models (VLMs) on structured diagram generation and editing tasks using a Diagram-as-Code approach with mxGraph XML. The benchmark features 1,449 diagrams from 6 domains and assesses models with metrics such as Execution Success Rate and Style Consistency Score. Experiments reveal that current state-of-the-art VLMs face significant challenges in maintaining structured fidelity and following instructions in these tasks.
Why it matters: VCG-Bench fills a key gap by providing a unified, structured evaluation framework for VLMs in professional diagrammatic applications, exposing current limitations in model capabilities.
Full story at: arXiv Computation and Language ↗