כתבה
arXiv cs.AI ·
לא כל הפרגמנטים שווים: פרגמנט-סקפולדינג מונחה-חקירה ללמידת-מודלים-רב-מודלי
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training
מחקר חדש מציע פרגמנט-סקפולדינג מונחה-חקירה ללמידת-מודלים-רב-מודלי. המחקר משתמש במודלי-למידה-גדולים-רב-מודליים (MLLMs) ומציע פתרון לבעיית הפרגמנטים השונים בלמידת-מודלים.
תקציר מקורי באנגליתarXiv:2609.15051v1 Announce Type: cross Abstract: Training prompts in online reinforcement learning (RL) differ substantially in how informative they are for the current policy: some are already saturated while others are too difficult to yield reliable learning signals, yet both receive equal rollout budget under standard training. We propose an exploration-guided prompt scaffolding framework that adapts the training prompt distribution dynamically throughout RL post-training of multimodal large language models (MLLMs). Central to our approach is the $\textit{Exploration Potential Score} (EPS)$, a lightweight rollout-based proxy for prompt utility derived from KL-regularized policy improvement theory, computable directly from on-policy rollout statistics without additional overhead. Rathe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית