יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מה נעבר מטקסט לראייה? חוקי הגדילה של יכולת ודינמיקות העברה ל-VLMs

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs
חברת המחקר הציגה תוכנה חדשה שמצפה את היכולת של VLMs מבוססת על ניסויי LLMs. התוכנה נקראת CDMScaling והיא ניתנת להורדה בגיטהאב.
תקציר מקורי באנגליתarXiv:2608.00013v3 Announce Type: replace-cross Abstract: Choosing the right large language model (LLM) backbone is the most consequential decision when building a vision-language model (VLM), yet it remains fundamentally unprincipled: compute-based scaling laws fail to generalize across model families, and no framework exists for directly predicting VLM performance before training begins. We propose the Capability-Driven Multimodal Scaling Law, the first cross-family framework that predicts VLM benchmark accuracy from directly observable textual capability. Given a low-dimensional capability score $S$ extracted from LLM textual benchmarks via PCA, we model VLM performance as a function of $S$, with a per-backbone transfer rate and an absorption rate that quantifies data-scaling efficiency
קרא במקור המקורי