← Back to brief
ResearchOfficialPreprintarXiv Computer Vision

Eddy-VL 1.9B: Structural Pruning and Layered Distillation for Edge-Deployable Multimodal Embedding

Researchers present Eddy-VL 1.9B, a compressed multimodal embedding model derived from Qwen3-VL-Embedding-2B, designed for offline, edge-deployable vision-language retrieval. The model employs structural pruning and layered knowledge distillation, reducing parameters by 9.5% while retaining 91.7% of the teacher model's performance on the MMEB-V2 benchmark. Eddy-VL is intended for air-gapped forensic and investigative scenarios where cloud APIs are inaccessible, and demonstrates reduced latency and strong performance on several compositional reasoning benchmarks.

Why it matters: This work provides a practical method for deploying efficient multimodal retrieval models on edge devices, supporting privacy-sensitive and low-latency applications where cloud access is not possible.

Full story at: arXiv Computer Vision