MissionBench Benchmark Shows Multimodal LLMs Struggle with Complex Aerial Tasks
A new arXiv preprint introduces MissionBench, a benchmark designed to evaluate multimodal large language models (MLLMs) on long-horizon, mission-level tasks in simulated aerial environments. The study finds that the best-performing MLLMs succeed on fewer than 35% of missions, while humans achieve 84.4%, revealing significant gaps in multi-step planning and adaptive reasoning. The results suggest that current general-purpose MLLMs are not yet capable of reliably handling complex embodied tasks without domain-specific training.
Why it matters: This work highlights a major limitation in current MLLMs, raising important questions about their readiness for real-world embodied applications and the risks of relying solely on scaling for improvement.
Full story at: arXiv Computation and Language ↗