כתבה
arXiv cs.AI ·
GeoPID: פירוק וניהול של מידע חזותי במודלי תקשורת רב-מודלי
GeoPID: Decomposing and Steering Visual Information in Vision-Language Models
במאמר זה, GeoPID, נציגים פרקטיקה חדשה לניתוח מידע חזותי במודלי תקשורת רב-מודלי. הם מציעים טכניקה של פירוק וניהול של מידע חזותי, שמאפשרת השפעה יעילה על יכולות הקריאה של המודל. המחברים מדגימים את יעילות הטכניקה באמצעות ניתוח של 22 מודלי תקשורת ו-14 מבחנים.
תקציר מקורי באנגליתarXiv:2610.08401v1 Announce Type: cross Abstract: While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over-rely on textual context. In this work, we propose \textsc{GeoPID}, a training-free framework that analyzes multimodal information within VLMs from a geometric perspective. \textsc{GeoPID} decomposes information into Redundant, Modality-Unique, and Synergistic components through the geometric relationships between visual and textual representation subspaces. Through an extensive analysis across 22 VLMs and 14 benchmarks, we confirm that correct predictions exhibit stronger vision-unique components when questions strongly require visual grounding. Building on this geometric analysis, we
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית