כתבה
arXiv cs.LG ·
Understanding Cross-Modal Contributions in Continual Vision-Language Models: A Theoretical Perspective
תקציר מקורי באנגליתarXiv:2606.14883v2 Announce Type: replace-cross Abstract: Continual vision-language models are commonly addressed through sequential fine-tuning; however, although this paradigm enables adaptation to new environments (tasks), it inherently emphasizes the contribution of previously learned environments (tasks) at the expense of the stability required to preserve previously acquired knowledge. While existing approaches have adequately studied continual learning and catastrophic forgetting in vision-language models (VLMs), the theoretical understanding of modality-specific contributions across a sequence of environments remains largely unexplored. In this paper, we present a new theoretical perspective to understand the cross-modal (vision-language) contributions to consecutive environments.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית