כתבה
arXiv cs.LG ·
Stable-MM-R1: יציבות תהליכי ניתוח רב-מודאליים
Stable-MM-R1: Anchoring Multimodal Reasoning Dynamics via Entropy-Guided Stratification
Stable-MM-R1 הוא כלי לייצוב תהליכי ניתוח רב-מודאליים. הוא משתמש בשיטות חדשות כדי לשפר את יציבות האימון ולמנוע קריסת אנטרופיה. החידוש משתמש ב-Potential-Aware Query Mining ו-Hybrid Stratified Replay.
תקציר מקורי באנגליתarXiv:2609.07148v1 Announce Type: new Abstract: While Reinforcement Learning (RL) effectively incentivizes reasoning in Large Language Models, current pipelines are hindered by training instability and rapid entropy collapse. These limitations often stem from "Rollout Silencing" and low-quality gradient signals in standard sampling procedures. In this work, we propose a robust, data-centric framework to stabilize RL training. We first introduce Potential-Aware Query Mining (PAQM), which filters data dynamically to focus on the "Distillation Zone"---samples with high potential for capability elicitation. Furthermore, we present Hybrid Stratified Replay (HSR), a novel mechanism that restructures batches by stratifying rollouts based on Path Entropy, a rollout-level confidence proxy, and outc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית