כתבה
arXiv cs.LG ·
לא כל ההנחות הן שוות: פרסום מודרך של הנחות ללמידה רפלקסיבית
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
הצעה לשיפור בלמידה רפלקסיבית: פרסום מודרך של הנחות ללמידה רפלקסיבית במודלי שפה גדולי-מימד. הצעה זו פותחת פרסום מודרך של הנחות ללמידה רפלקסיבית, שמאפשרת למודלים ללמוד בצורה יעילה יותר. הצעה זו יכולה לשפר את יעילות הלמידה רפלקסיבית ולהפחית את הזמן הדרוש ללמידה.
תקציר מקורי באנגליתarXiv:2609.15051v1 Announce Type: new Abstract: Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard training. We propose an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs). Central to our approach is the $\textit{Exploration Potential Score} (EPS)$, a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory, computable directly from on-policy rollout statistics without additional overhead. Rather
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית