יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מדידת הפער בין הבנה והפקה במודלים מרוכזים מרובדיים

Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
במאמר זה, נחקור את הפער בין הבנה והפקה במודלים מרוכזים מרובדיים. נציג כלי חדש לבדיקה, GapEval, ונבחן את המודלים השונים.
תקציר מקורי באנגליתarXiv:2602.02140v2 Announce Type: replace Abstract: Recent advances in unified multimodal models (UMM) have demonstrated remarkable progress in both understanding and generation tasks. However, whether these two capabilities are genuinely aligned and integrated within a single model remains unclear. To investigate this question, we introduce GapEval, a bidirectional benchmark designed to quantify the gap between understanding and generation capabilities, and quantitatively measure the cognitive coherence of the two "unified" directions. Each question can be answered in both modalities (image and text), enabling a symmetric evaluation of a model's bidirectional inference capability and cross-modal consistency. Experiments reveal a persistent gap between the two directions across a wide rang
קרא במקור המקורי