יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Inverse-LLaVA: חשיבה מחדש של התאמה מודלית באמצעות מיפוי טקסט-לתמונה

Inverse-LLaVA: Rethinking Multimodal Alignment via Text-to-Vision Mapping
המאמר Inverse-LLaVA חושף חידוש בתחום התאמה מודלית, כאשר הוא מציע פתרון חדש לבעיית ההתאמה בין דגמי תמונה ולשון. החידוש נקרא Inverse-LLaVA והוא מבוסס על טכניקת מיפוי טקסט-לתמונה. המאמר כולל ניסויים שונים וביניהם ניתוח של התאמה והשפעת הפתרון על תחומים שונים.
תקציר מקורי באנגליתarXiv:2508.12466v3 Announce Type: replace-cross Abstract: Connecting pretrained vision and language models usually involves projecting image features into the language model's input space. Inverse-LLaVA reverses this mapping within decoder attention: language states are projected to the visual feature dimension, and modality-specific maps produce residual query, key, and value updates. Fusion and low-rank adaptation (LoRA) learn jointly from 665K visual instructions, with frozen backbones and no separate alignment stage. Across nine primary benchmark evaluations, the final 7B model approaches two-stage LLaVA-1.5 on several tasks. It scores 78.45% on VQAv2 versus 79.13% for official LLaVA-LoRA, and 50.96% versus 48.56% on VizWiz; TextVQA is lower at 56.96% versus 58.47%. Controlled studies
קרא במקור המקורי