← Back to arXiv Computer Vision

arXiv Computer Vision briefings

ResearchOfficialarXiv Computer Vision

Group-Contrastive Forward-Forward Algorithm Yields Hierarchical Monosemantic Neurons

Researchers introduce the Group-Contrastive Forward-Forward (GCFF) algorithm, a biologically inspired training method that produces monosemantic neurons organized in hierarchies of increasing abstraction. Unlike sparse autoencoders, GCFF captures non-linear concepts without relying on sparsity constraints and achieves state-of-the-art performance among forward-forward algorithms on image classification benchmarks.

Why it matters: This work suggests a new approach to mechanistic interpretability by showing that monosemanticity can emerge from local, layer-wise learning rules, potentially enabling more interpretable neural networks.

ResearchOfficialarXiv Computer Vision

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

Researchers introduce LookME, a framework that enables lookup-based enhancement for multimodal embeddings in vision-language models (VLMs). LookME employs a hierarchical two-level lookup method and a sparse injection strategy to efficiently retrieve and inject relevant multimodal embeddings, supporting partitioned storage and on-demand loading. Experimental results indicate that LookME outperforms text-only PLE-style methods on multiple visual benchmarks, demonstrating improved efficiency and performance.

Why it matters: LookME provides a memory-efficient approach to scaling vision-language models, potentially enabling their deployment in resource-constrained environments.

ResearchOfficialarXiv Computer Vision

Apple-PI Benchmark Tests Video Models' Grasp of Physical Laws

Researchers introduce Apple-PI, a benchmark designed to evaluate video generation models based on their adherence to physical laws, rather than just output plausibility. The benchmark features 400 videos covering classical mechanics tasks and employs a three-stage protocol to diagnose where models fail in the reasoning process. Testing 11 models revealed that the best-performing model achieved a score of only 0.473, highlighting that current video models are not yet reliable law-grounded world simulators.

Why it matters: Apple-PI offers a new diagnostic tool to guide the development of video models toward genuine physical understanding, which is essential for applications that require accurate world simulation.

ResearchOfficialarXiv Computer Vision

PhysAgent: Reflective Agentic Framework for Physically Plausible Video Generation

PhysAgent is a reflective agentic framework designed to improve physics-grounded video generation by iteratively generating, simulating, verifying, and repairing physical programs. It introduces a set of physics-control APIs to enable more stable and complex motion behaviors. Experiments indicate that PhysAgent produces more physically plausible videos and achieves better prompt alignment compared to prior approaches.

Why it matters: This work advances the reliability and complexity of physics-based video synthesis, addressing the challenge of accurately translating user intent into physically plausible video outputs.

ResearchOfficialarXiv Computer Vision

CRISP: Pre-LLM Text-Driven Visual Token Pruning for Efficient LVLM Inference

Researchers introduce CRISP, a two-stage framework for pruning visual tokens before input to large vision-language models (LVLMs), guided by text to retain instruction-relevant evidence and scene context. Experiments on LLaVA-1.5 and LLaVA-NeXT show that CRISP can maintain up to 99.5% of original accuracy while reducing inference cost and latency by more than 2x. This approach addresses the inefficiency of processing large numbers of visual tokens in LVLMs, particularly for resource-constrained environments.

Why it matters: CRISP enables more efficient deployment of LVLMs by significantly reducing inference overhead without major accuracy loss, making advanced multimodal AI more accessible in limited-resource settings.

ResearchOfficialarXiv Computer Vision

3D FaceShell: Attribute Transfer in 3D Face Avatars as a VLM Defense Mechanism

Researchers introduce 3D FaceShell, a framework that applies subtle, learnable perturbations to 3D face avatars to mislead vision-language models (VLMs) from accurately inferring sensitive attributes, while preserving the avatar's visual fidelity and identity. The method uses a Gaussian shell optimized via multi-view embedding alignment to redirect VLM-based attribute inference. Experiments on celebrity face avatars and multiple black-box VLMs show that 3D FaceShell increases attribute injection and mismatch rates without compromising human-recognizable appearance.

Why it matters: This work proposes a novel defense against privacy risks posed by VLMs extracting sensitive information from 3D face avatars, operating directly on 3D representations rather than 2D images.

ResearchOfficialarXiv Computer Vision

Brain-Encoding Model's Predicted Responses Outperform Visual Backbone for Video Memorability on One Dataset, Not Another

A preprint study evaluates the TRIBE v2 brain-encoding model's predicted fMRI responses as features for forecasting video memorability. On the VideoMem dataset, these brain-encoded features outperform the model's own visual backbone, while on Memento10k, the backbone performs better. The dataset-specific advantage is robust to controls for sample size and feature compression, and a vision-orthogonal memorability signal is localized to the ventral occipito-temporal cortex.

Why it matters: The findings suggest that brain-encoding models can capture human-relevant behavioral signals missed by standard vision models, but their utility depends on the dataset, underscoring the importance of diverse benchmark evaluation.

ResearchOfficialarXiv Computer Vision

FedDP-PALD: Privacy-Preserving Federated Latent Diffusion for Medical Data Synthesis

Researchers introduce FedDP-PALD, a federated latent diffusion framework designed to generate synthetic medical images and ECG signals with formal differential privacy guarantees. The approach uses prototype aggregation with calibrated noise to protect against membership inference attacks while maintaining diagnostic utility, achieving F1 scores and AUROC values close to those obtained with real data. The method is evaluated on multiple medical datasets and demonstrates strong privacy protection with minimal loss in predictive performance.

Why it matters: This work offers a significant advance in privacy-preserving synthetic data generation for medical applications, enabling collaborative model training across institutions without compromising patient privacy.

ResearchOfficialarXiv Computer Vision

Model Merging for Medical LVLMs: A Benchmark and a Winner-Take-All Approach

Researchers introduce MergeMedBench, the first comprehensive benchmark for merging medical vision-language models (LVLMs), covering eight imaging modalities and diverse clinical tasks. They propose a winner-take-all merging method that retains only the most dominant parameters from expert models, avoiding the information dilution seen in averaging or alignment-based strategies. This hyperparameter-free approach consistently outperforms existing merging methods in their evaluations.

Why it matters: This work offers a practical and effective solution for consolidating multiple specialized medical LVLMs, potentially reducing deployment costs and complexity while maintaining strong performance.

ResearchOfficialarXiv Computer Vision

ImprovedVBGS: Real-time Continual Variational Bayes Gaussian Splatting

ImprovedVBGS is a new framework for real-time, on-the-fly 3D reconstruction, designed for applications in robotics and autonomous navigation. It achieves a 1680x speed-up over previous Variational Bayes Gaussian Splatting (VBGS) methods, reducing per-frame latency from approximately 84 seconds to 0.050 seconds on an RTX 3070 Ti, while maintaining reconstruction quality. The acceleration is accomplished through spatially truncated variational inference and improved reassignment strategies.

Why it matters: This advance enables practical, real-time continual 3D reconstruction for robotics and autonomous systems operating under strict latency and memory constraints.

ResearchOfficialarXiv Computer Vision

Debiasing Text-to-Image Evaluation via Implicit Cultural Alignment Reward Modeling

Researchers present an Implicit Cultural Alignment Reward Model based on a 4.2B-parameter multimodal large language model (MLLM) to assess cultural authenticity in text-to-image (T2I) outputs. The model achieves 80.54% pairwise accuracy on the CulturalFrames benchmark and processes each evaluation in 0.21 seconds, representing a 10x speedup over standard VQA-based evaluators. The approach outperforms existing vision-language metrics and MLLM-based evaluators in capturing culturally salient details.

Why it matters: This work offers a more efficient and culturally sensitive method for evaluating generative AI outputs, addressing biases often overlooked by current metrics.

ResearchOfficialarXiv Computer Vision

MoD-VLLM: Modularized Dynamic-Granularity Video LLM for Multi-Event Long Video Understanding

A new framework called MoD-VLLM is proposed for multi-event long video understanding. The approach iteratively localizes question-relevant video segments using a grounding module and a reflection module with dynamic granularity encoding. It employs a reinforcement learning strategy to jointly optimize grounding policies and visual representations. MoD-VLLM demonstrates significant improvements over state-of-the-art baselines on several benchmarks, including the newly introduced MEventBench.

Why it matters: This work introduces a modular and self-reflective approach that addresses the challenge of understanding multiple events in long videos, a key limitation of current Video LLMs.

ResearchOfficialarXiv Computer Vision

Dataset-Origin Signatures and Shortcut Learning in Screening Mammography AI: A Cross-Dataset Case Study

A new study demonstrates that adding biopsy-confirmed, abnormal-enriched cases from external datasets to screening mammography AI training can significantly reduce model performance, with AUC-ROC dropping from 0.737 to as low as 0.620. The performance decline is linked to persistent dataset-specific characteristics that dominate learned representations, even after consistent preprocessing. The research underscores that naive pooling of heterogeneous mammography datasets introduces domain shifts that can outweigh the benefits of increased positive cases.

Why it matters: This work highlights that hidden dataset biases can undermine AI reliability in medical imaging, emphasizing the need for domain-aware strategies when combining data from different sources.

ResearchOfficialarXiv Computer Vision

Training-Free Frame Selection for Long Videos Using Attention-Based MLLM Selectors

Researchers introduce DAFS, a training-free method for selecting frames in long videos by leveraging cross-modal attention from multimodal large language models (MLLMs). DAFS identifies query-relevant frames without requiring autoregressive generation or additional training, and formulates frame selection as a discrete optimization problem. The approach improves performance over uniform sampling by up to 6.4 points on Video-MME under a 32-frame budget and generalizes across different model backbones and tasks.

Why it matters: This work offers a practical advance for efficient, query-aware frame selection in long-video understanding, potentially broadening the applicability of MLLMs to real-world video analysis tasks.

ResearchOfficialarXiv Computer Vision

DiTango: Cost-Effective Parallel Diffusion Generation with Selective Attention State Reuse

Researchers introduce DiTango, a parallel framework for Diffusion Transformers that selectively reuses attention states to reduce communication overhead in multi-node environments. DiTango achieves up to 1.9x end-to-end and 3.2x attention speedup, with near-linear scaling, while maintaining generation quality comparable to state-of-the-art methods. The framework uses an anchor-guided state selection planner and a runtime for efficient state-centric operations.

Why it matters: DiTango addresses a key scalability bottleneck in diffusion model inference, enabling faster and more cost-effective high-resolution content generation in distributed settings.

ResearchOfficialarXiv Computer Vision

SlotMem: Character-Addressable Internal Memory for Narrative Long Video Generation

Researchers introduce SlotMem, a character-addressable internal memory framework designed for multi-character narrative long video generation. SlotMem employs a Character-Semantic Probe and Memory Encoder to compress visual tokens into role-specific slot memories, enabling more precise tracking of character identities. Experimental results on multiple benchmarks demonstrate that SlotMem improves long-range character consistency compared to existing methods, while maintaining similar video quality.

Why it matters: Maintaining consistent character identities across scene transitions is a major challenge in narrative video generation, and SlotMem offers a novel solution that advances this capability.

ResearchOfficialarXiv Computer Vision

Rethinking the Readout: Unlocking Video Backbones for AI-Generated Video Detection

A new method, Velocity Gated Patch Velocity Profiling (V-PVP), improves detection of AI-generated videos by replacing the standard global readout in video backbone models with a lightweight module that preserves local temporal artifacts. V-PVP achieves a 95.28 AUC on the AIGVDBench benchmark using a frozen backbone and adds only about 0.5 million parameters. The approach is plug-and-play and consistently boosts performance across various video backbones.

Why it matters: This work offers a practical and effective solution to enhance the detection of AI-generated videos, addressing a key challenge in deepfake detection.

ResearchOfficialarXiv Computer Vision

Reasoning-Guided Part-Level Visual Grounding via Reinforcement Learning

A new method, Object-Part Hierarchical Reflective Grounding (OP-HRG), is introduced to address the challenge of part-level visual grounding in multimodal large language models. OP-HRG employs a coarse-to-fine reasoning strategy, first localizing the parent object and then the specific part, with a reflective self-check mechanism. Trained with a part-aware reinforcement learning framework, the approach achieves state-of-the-art results on several part grounding benchmarks, outperforming larger existing models.

Why it matters: This work advances fine-grained visual understanding in multimodal models, enabling more accurate part-level grounding for applications such as robotics and image editing.

ResearchOfficialarXiv Computer Vision

AE-UAV: First Air-to-Air Event-Based UAV Tracking Benchmark and Real-Time CPU Tracker

Researchers have introduced AE-UAV, the first airborne-captured event camera dataset specifically for air-to-air UAV tracking, featuring 178 flight sequences with detailed annotations. They also present FSFT, a lightweight, training-free tracker that achieves 420 FPS on CPU-only hardware and retains 93.97% of the accuracy of state-of-the-art GPU-based methods. This approach offers a 5.32-fold speedup and demonstrates strong generalization in temporal resolution.

Why it matters: This work enables efficient, real-time UAV tracking on resource-constrained platforms, addressing a major challenge in airborne remote sensing.

ResearchOfficialarXiv Computer Vision

Ego Scene Augmentation Framework Improves Spatial Perception in Multimodal LLMs

Researchers have introduced Ego Scene Augmentation (ESA), a framework designed to enhance egocentric spatial perception in multimodal large language models (MLLMs) by leveraging an Ego-element Graph. ESA delivers notable improvements on the EgoTextVQA benchmark, achieving 8.14% and 8.72% gains in indoor and outdoor settings, respectively, and demonstrates particularly strong results in the shopping subset of the indoor setting.

Why it matters: Improving spatial reasoning in egocentric scenes addresses a key challenge for MLLMs, which is essential for advancing real-world interaction capabilities.