Multimodal AI news — Page 4

AI systems that understand or generate combinations of text, images, audio, video, and other types of data.

ResearchOfficialarXiv Computation and Language

VCG-Bench: A Unified Benchmark for Structured Diagram Generation and Editing with Vision-Language Models

Researchers introduce VCG-Bench, a benchmark designed to evaluate vision-language models (VLMs) on structured diagram generation and editing tasks using a Diagram-as-Code approach with mxGraph XML. The benchmark features 1,449 diagrams from 6 domains and assesses models with metrics such as Execution Success Rate and Style Consistency Score. Experiments reveal that current state-of-the-art VLMs face significant challenges in maintaining structured fidelity and following instructions in these tasks.

Why it matters: VCG-Bench fills a key gap by providing a unified, structured evaluation framework for VLMs in professional diagrammatic applications, exposing current limitations in model capabilities.

ResearchOfficialarXiv Computation and Language

ActiveVision Benchmark Shows MLLMs Struggle with Active Visual Observation

A new preprint introduces ActiveVision, a benchmark designed to test whether multimodal large language models (MLLMs) can perform active visual observation—redirecting their 'gaze' based on intermediate reasoning, rather than relying on static images. Leading models such as GPT-5.5 and Claude Fable 5 scored only 10.6% and 3.5% respectively, compared to a human average of 96.1%. The findings suggest that current MLLMs lack robust active visual perception, even when allowed to write and execute their own vision code.

Why it matters: This work reveals a fundamental limitation in current MLLMs, highlighting the need for new architectures that integrate perception and reasoning in a closed loop.

ResearchOfficialarXiv Cryptography and Security

Black-box Visual Attacks on Long-term Memory in Multimodal AI Agents

Researchers introduce Lucid, a black-box adversarial framework that targets multimodal AI agents by crafting imperceptible perturbations to images, compromising their long-term memory pipelines. Lucid enables two attack modes—memory poisoning and memory injection—achieving 61.6% and 58.4% attack success rates, respectively, across five black-box memory architectures, including commercial systems. The attacks require no access to the target model or text channel, operating solely through manipulated visual inputs.

Why it matters: This work exposes a critical vulnerability in multimodal AI agents' reliance on visual data for persistent memory, highlighting the risk of adversarial manipulation through images alone.

ResearchOfficialarXiv AI/ML

AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning

AV-JEPA is a multimodal extension of LeJEPA for audio-visual self-supervised learning, employing an early-fusion Vision Transformer and modality dropout. The architecture achieves competitive results on VGGSound (57.1% top-1) and AudioSet (32.7 mAP), and enables zero-shot audio-video retrieval without requiring decoders or contrastive negatives.

Why it matters: AV-JEPA demonstrates a streamlined approach to multimodal self-supervised learning, achieving strong performance on audio-visual tasks with a simplified architecture.

ResearchOfficialarXiv AI/ML

SLAPBench: First Benchmark for MLLM-Based Four-Finger SLAP Fingerprint Verification

Researchers introduce SLAPBench, the first benchmark for evaluating multimodal large language models (MLLMs) on four-finger SLAP fingerprint verification, using data from NIST SD302b with 7,832 pairs. The study finds that prompting strategies determine whether verification collapses, while model capability affects discrimination performance. Claude Opus 4.8 achieves the best binary verification result (FAR=20.2%), and Qwen3-VL-8B achieves perfect separation (AUC=1.000) under similarity scoring, though this may reflect dataset artifacts rather than true capability.

Why it matters: This work establishes the first SLAP-specific MLLM baseline and reveals how prompting and model capability interact in high-stakes biometric verification tasks.

Policy & SafetyOfficialarXiv AI/ML

FLINT: Fingerprinting Federated Learning Architectures from 5G PHY-Layer Side Channels

Researchers present FLINT, a black-box framework capable of inferring federated learning model architecture families (such as CNNs, RNNs, and Transformers) by analyzing only 5G physical-layer side-channel information. FLINT operates without access to packet-level data, instead leveraging scheduling metadata from the 5G Physical Downlink Control Channel (PDCCH) to identify temporal patterns linked to specific model architectures. In over-the-air experiments, FLINT achieves a macro F1-score of 0.930 for architecture-family classification, demonstrating a new class of side-channel leakage in federated learning over 5G networks.

Why it matters: This work reveals a previously unrecognized security vulnerability in federated learning over 5G, showing that model architectures can be fingerprinted via physical-layer side channels, potentially enabling targeted attacks.

ResearchOfficialarXiv AI/ML

Prompt Echoing Resolves Question-First Paradox in Vision-Language Models

Researchers have identified a 'question-first paradox' in vision-language models (VLMs), where placing the question before the image in prompts—though intuitive—leads to worse performance than placing the image first. Through analysis, they attribute this to a trade-off between steering perception and maintaining question accessibility at answer time. They propose a training-free solution called 'question echoing,' which involves restating the question both before and after the image in the prompt. This method closes the performance gap and improves accuracy by up to 19 points on several benchmarks, without requiring any model retraining or architectural changes.

Why it matters: This finding offers a simple, immediate way to boost VLM performance through prompt design alone, benefiting users and developers without additional computational cost.

ModelsOfficialarXiv AI/ML

S1-Omni: A Unified Multimodal Reasoning Model for Scientific Understanding, Prediction, and Generation

S1-Omni is a unified multimodal reasoning model designed for a wide range of scientific tasks, including property prediction, spectrum-to-molecular generation, and protein structure prediction. The model consolidates capabilities that were previously fragmented across domain-specific models, mapping diverse scientific data and natural-language instructions into a shared representation space. According to its preprint, S1-Omni outperforms GPT-5.5 and Gemini-3.1-Pro on most scientific benchmarks and matches or surpasses specialized models on several tasks.

Why it matters: S1-Omni represents a significant step toward unified AI models for science, potentially streamlining research workflows and reducing the need for multiple specialized systems.

ResearchOfficialarXiv AI/ML

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

DrawingVQA introduces the first benchmark specifically designed to evaluate multimodal large language models on real-world construction drawings. The dataset includes 33 authentic construction drawings and 92 expert-curated question-answer pairs, covering three levels of reasoning complexity. Evaluations show a significant performance gap between current models and human experts, especially at more advanced reasoning levels.

Why it matters: This benchmark provides a crucial resource for advancing AI capabilities in domain-specific multimodal reasoning, with direct relevance to engineering and construction workflows.

ResearchOfficialarXiv AI/ML

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

SeerGuard is a safety framework for mobile GUI agents that introduces pre-execution instruction-level screening and action-level risk assessment. It employs a safety-augmented world model (SAWM) to predict the outcomes of agent actions and assess potential risks before execution. Experimental results show that SeerGuard improves safety-utility scores and reduces risk-cost scores across various agents, demonstrating effective generalization.

Why it matters: SeerGuard enables proactive risk assessment for mobile GUI agents, addressing a key safety challenge by helping prevent irreversible errors before they occur.

ResearchOfficialarXiv AI/ML

MAR-12: Multi-Angle Reasoning Framework for Detecting and Explaining Harmful Humor in Memes

Researchers have introduced MAR-12, a novel framework that leverages Vision Language Models to detect and explain harmful humor in memes by analyzing twelve structured perspectives based on humor and hate theories. MAR-12 achieves up to 80.3% accuracy for humor detection and 75.9% for hate detection on benchmark datasets, outperforming previous state-of-the-art methods. The system generates transparent, context-grounded explanations for its decisions, particularly in cases where humor and hate coexist. Human and GPT-4-based evaluations confirm the coherence and persuasiveness of its explanations.

Why it matters: This work advances explainable AI for multimodal content moderation, addressing the challenge of interpreting memes where humor and harmful intent overlap.

ModelsReportedThe Decoder

Alibaba unveils Qwen 3.8, a 2.4-trillion-parameter multimodal model

Alibaba has introduced Qwen 3.8, a multimodal AI model with 2.4 trillion parameters. According to the Qwen team, it rivals leading models and is second only to Fable 5. A preview of the model is currently available.

Why it matters: This release highlights the growing competition in large-scale open-weight multimodal AI models.

ResearchOfficialApple Machine Learning Research

Apple Introduces VICIS: Benchmarking Visual Concept Inference from Image Sets

Apple Machine Learning Research has introduced Visual Concept Inference from Sets (VICIS), a new task designed to evaluate whether vision-language models (VLMs) can infer shared concepts from small sets of example images and apply them to new queries. The research finds that current state-of-the-art VLMs perform poorly on this benchmark, revealing a significant limitation in their visual reasoning abilities.

Why it matters: This benchmark highlights a key gap in vision-language models' ability to learn and generalize visual concepts from limited visual context, which is important for advancing few-shot learning in AI.

ResearchOfficialarXiv Software Engineering

FirmPilot: Multi-Agent Framework Significantly Improves IoT Firmware Rehosting

FirmPilot is an evidence-guided multi-agent framework designed to enhance IoT firmware rehosting by iteratively recovering boot, state, and network artifacts. In evaluations on the LFwC corpus, FirmPilot increased web-service reachability from 25.49% to 52.39% and network reachability from 39.30% to 71.93% compared to the prior FirmAE system. The framework also raised the average number of detected services per firmware and enabled downstream analysis workflows such as RouterSploit interaction and protocol-aware fuzzing.

Why it matters: This work represents a substantial advance in automated IoT firmware analysis, improving the reliability and scalability of dynamic analysis for security testing and vulnerability discovery.

ResearchOfficialarXiv Robotics

Open-AoE: Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning

Open-AoE is an open, community-oriented dataset and toolchain for egocentric manipulation, featuring approximately 2,000 hours of video collected by over 500 contributors using more than 400 smartphones. The dataset includes structured annotations such as text labels, MANO-based hand poses, camera trajectories, and temporally localized atomic actions. It also provides a full pipeline for data processing, including temporal action segmentation, semantic annotation, hand and camera trajectory reconstruction, as well as downstream tools for visualization, cross-embodiment retargeting, and model training. This resource is designed to facilitate embodied AI research and human-to-robot transfer by lowering barriers to data contribution and reuse.

Why it matters: Open-AoE provides a large-scale, practical infrastructure for embodied AI research, potentially accelerating advances in manipulation tasks and world modeling.

ResearchOfficialarXiv Computer Vision

Ego Scene Augmentation Framework Improves Spatial Perception in Multimodal LLMs

Researchers have introduced Ego Scene Augmentation (ESA), a framework designed to enhance egocentric spatial perception in multimodal large language models (MLLMs) by leveraging an Ego-element Graph. ESA delivers notable improvements on the EgoTextVQA benchmark, achieving 8.14% and 8.72% gains in indoor and outdoor settings, respectively, and demonstrates particularly strong results in the shopping subset of the indoor setting.

Why it matters: Improving spatial reasoning in egocentric scenes addresses a key challenge for MLLMs, which is essential for advancing real-world interaction capabilities.

ResearchOfficialarXiv Computer Vision

GeoDetect: Geometric Adversarial Detection for Vision-Language Pre-trained Models

Researchers have introduced GeoDetect, a method that utilizes the geometric properties of embedding spaces in vision-language pre-trained models (VLPs) to detect adversarial examples. By analyzing the anisotropic structure of VLP embeddings, they found that adversarial examples tend to have greater distances to random points compared to clean examples. GeoDetect leverages this property to reliably identify adversarial attacks across various VLP architectures and threat scenarios, including both unimodal and multimodal attacks.

Why it matters: This work offers a robust and practical approach to detecting adversarial attacks in vision-language models, enhancing their safety and reliability.

ResearchOfficialarXiv Computer Vision

MonteRET: AI Agent Enhances Chest CT Report Generation with Multi-granularity Knowledge Retrieval

Researchers have introduced MonteRET, a region-aware retrieval-enhanced framework for automated chest CT report generation. The system combines global and region-level CT features, retrieves clinically relevant knowledge based on predicted conditions and anatomical regions, and uses an AI agent to refine initial reports. MonteRET demonstrated improved report quality, semantic similarity, and clinical efficacy compared to baselines and state-of-the-art methods on both public and external datasets, with human expert evaluations favoring its outputs.

Why it matters: This work shows that integrating multi-granularity knowledge retrieval and vision-language alignment can significantly enhance the clinical accuracy and completeness of automated radiology report generation.

ResearchOfficialarXiv Computer Vision

SD-MAR: Multi-image Analytical Reasoning via Synthetic Data and Reinforcement Learning

Researchers present SD-MAR, a framework for training and evaluating vision-language models on multi-image analytical reasoning tasks, such as change detection and quantitative comparison. By leveraging synthetic data and a reinforcement learning method called GRPO-lite with Backward Discounted Allocation, they report up to 36.95% accuracy improvement on in-domain benchmarks. Notably, Qwen2.5-VL-7B outperforms GPT-4.1 on the SD-MAR benchmark, and out-of-domain generalization is maintained or improved on several standard benchmarks.

Why it matters: This work advances multimodal AI by enabling models to reason analytically across multiple images, a capability important for real-world tasks involving visual comparison and inference.

ResearchOfficialarXiv Computer Vision

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents

AdaTurn is a framework for active visual perception agents that enables them to adapt to varying rollout budgets by conditioning on the allowed number of turns and explicitly training for boundary behavior. Its Forced-Answer DAPO component turns over-budget events into trainable final-decision steps, allowing agents to synthesize answers even when further actions are not possible. The approach significantly improves low-budget accuracy and generalizes across different agent backbones and multimodal benchmarks.

Why it matters: AdaTurn addresses a key deployment challenge by enabling active visual agents to provide valid answers under varying and tight turn limits, improving their practical usability.