כתבה
arXiv cs.AI ·
העברת עדכוני RL באופן סלקטיבי לתפיסה ראייתית
Selective Transfer of RL Updates for Visual Reasoning
במאמר זה, המחברים מציגים שיטה חדשה להעברת עדכוני RL לתפיסה ראייתית. השיטה, הקרויה Selective-RL, מבוססת על העברת עדכוני RL באופן סלקטיבי, כדי לשמר את המידע החשוב ביותר. המחברים מדגימים את השיטה באמצעות ניסויים, ומציגים תוצאות טובות יותר מאשר שיטות קודמות.
תקציר מקורי באנגליתarXiv:2610.08659v1 Announce Type: cross Abstract: Model merging provides a training-free way to transfer reasoning capabilities from language models to vision-language models (VLMs), but endpoint-based transfer can conflate pre-existing model differences with changes acquired during reasoning post-training. We instead formulate capability transfer around the training-stage update, isolating the parameter changes induced by reinforcement learning (RL). Yet transferring this update in full remains suboptimal: we find that its components differ substantially in cross-model transferability, with dominant directions transferring more effectively than the complete update. Based on this finding, we introduce Selective-RL, which isolates the RL-stage update, retains its dominant matrix-wise direct
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית