Dual Adversarial Fine-tuning Improves Robustness of Large Vision-Language Models Across Tasks
A new dual adversarial fine-tuning framework has been proposed to enhance the robustness of large vision-language models (LVLMs) against adversarial attacks. By jointly optimizing visual and semantic supervision signals, the method generalizes across multiple tasks—including zero-shot classification, image captioning, and visual question answering—without requiring task-specific retraining. Experimental results indicate that this approach outperforms existing state-of-the-art defense methods in adversarial robustness evaluations.
Why it matters: This work offers a generalizable defense mechanism that addresses the vulnerability of LVLMs to adversarial attacks across diverse multimodal tasks.
Full story at: arXiv Computer Vision ↗