כתבה
arXiv cs.CL ·
FORUM: פתרון תמונות קפואות באמצעות הסכמה של מודלים להתייחסות ויזואלית
FORUM: Frozen Outputs Reconciled Using Model Agreement for Visual Grounding
מודלי שפה גדולים קפואים ומודלים מודאליים פותרים בהצלחה התייחסות תקינה למושגים ויזואליים באמצעות קריאת פקודה אחת.
תקציר מקורי באנגליתarXiv:2609.37488v1 Announce Type: cross Abstract: Frozen multimodal large language models (MLLMs) now solve standard referring expression comprehension with a single prompted call, yet on adversarial benchmarks with same-category distractors and negation, even the largest models are confidently wrong, and resampling repeats the error. Models built from different data and architectures rarely fall for the same confounder, so their agreement is a strong label-free signal of the correct target. We present FORUM, a training-free test-time fusion of frozen MLLMs guided by two fixed geometric rules: agreement-based selection keeps the region supported by the most distinct models, and medoid localization returns an actual member box instead of a coordinate average, so one loose prediction cannot
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית