יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הפחתת חדירת מידע בין תמונות במודלי תפיסה רב-תפיסה

Mitigating Cross-Image Information Leakage in Multi-Image Understanding with Large Vision-Language Models
במאמר זה, המחברים מציגים פתרון לבעיה של חדירת מידע בין תמונות במודלי תפיסה רב-תפיסה. הם מציגים את FOCUS, תהליך המסתפק במודלים קיימים ומשפר את הביצועים במשימות תפיסה רב-תפיסה.
תקציר מקורי באנגליתarXiv:2508.13744v3 Announce Type: replace-cross Abstract: Large Vision-Language Models (LVLMs) exhibit strong performance on single-image tasks. However, their performance degrades significantly when handling multi-image inputs. While this degradation has been observed in prior work, its nature remains poorly understood. We empirically observe visual elements from different images become entangled in the model's representations and responses. We refer to this phenomenon as cross-image information leakage. To address this issue, we propose FOCUS, a training-free and architecture-agnostic method. FOCUS masks all but one image with random noise, guiding the model to focus on the single clean image. This process is applied across the target images to obtain logits under partially masked contex
קרא במקור המקורי