כתבה
arXiv cs.LG ·
Where-OPD: גידור חללי מודלי ML עם תמונות סינתטיות
Where-OPD: Spatially Guided On-Policy Self-Distillation of MLLMs with Synthetic Scenes
Where-OPD הוא שיטה לשיפור יכולות ההבנה של מודלי ML גדולי-הרשת, על ידי שימוש בתמונות סינתטיות והנחיית חללי. השיטה נועדה לשפר את היכולות של המודלים להבין תמונות ולזהות חלקים רלוונטיים.
תקציר מקורי באנגליתarXiv:2610.02117v1 Announce Type: cross Abstract: On-policy self-distillation has recently emerged as an effective approach for improving language-model reasoning by supervising students with a frozen or EMA version of themselves that receives privileged information. Its application to multimodal large language models (MLLMs), however, remains largely unexplored. Recent approaches use privileged visual information, such as image crops corresponding to a question, to improve fine-grained perception, but their gains are confined to tasks that benefit from such visual zooming and require either human-annotated grounding data or external teacher models. We introduce a different form of on-policy self-distillation for MLLMs that provides the teacher with textual, spatially grounded guidance ide
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית