כתבה
arXiv cs.AI ·
גבולות הקנה במודלים ויז'ואליים-לשוניים
Not All Error Yields to Scale: Where Scaling Stops in Vision-Language Inference
מודלים ויז'ואליים-לשוניים מתמודדים עם פשרות בין עיבוד מידע ויז'ואלי לשימוש במודל שפה גדול יותר. מחקר זה מציג את חוק הפרדה, המתאר כיצד ביצועי המודל משתנים עם גודל המודל וכמות הטוקנים הוויז'ואליים. המחקר בדק 26 מודלים שונים, כולל QwenVL, ומצא כי הרווחים מגודל מודל גדול יותר דומים, אך הרווחים מטוקנים ויז'ואליים רבים יותר שונים.
תקציר מקורי באנגליתarXiv:2610.01640v1 Announce Type: cross Abstract: Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which combination of backbone size and input resolution to deploy, especially in high-resolution deployments. To address this gap, we propose the Separable Law that describes how VLM performance changes with language backbone size and visual token count. We fit the law to measurements from 26 InternVL and QwenVL models, with language backbone sizes from 1B to 72B, on four high-resolution benchmarks with image sizes from 224 pixels to 8K. We find that the questions responding to scaling can be predicted from the skill they
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית