AV-JEPA: Extending LeJEPA to Audio-Visual Self-Supervised Learning
AV-JEPA is a multimodal extension of LeJEPA for audio-visual self-supervised learning, employing an early-fusion Vision Transformer and modality dropout. The architecture achieves competitive results on VGGSound (57.1% top-1) and AudioSet (32.7 mAP), and enables zero-shot audio-video retrieval without requiring decoders or contrastive negatives.
Why it matters: AV-JEPA demonstrates a streamlined approach to multimodal self-supervised learning, achieving strong performance on audio-visual tasks with a simplified architecture.
Full story at: arXiv AI/ML ↗