יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אופטימיזציה מונחה רציונל: למידה להסביר עם ספסל סקפולדינג רציונלי

Rationale-Guided Policy Optimization: Learning to Reason with Adaptive Rationale Scaffolding
אופטימיזציה מונחה רציונל: פרקטיקה חדשה לשיפור יכולות ההסבר של מודלי LLM. המאמר מציג פרקטיקה חדשה לשיפור יכולות ההסבר של מודלי LLM, המשתמשת בספסל סקפולדינג רציונלי. הפרקטיקה, הקרויה RGPO, מאפשרת למודלים ללמוד להסביר עם סקפולדינג רציונלי, ולשפר את יכולות ההסבר שלהם.
תקציר מקורי באנגליתarXiv:2610.07342v1 Announce Type: new Abstract: On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is often limited by reward sparsity: when a model fails to discover correct trajectories for difficult problems, the optimization process receives little useful signal and may stagnate. Existing approaches mitigate this issue by incorporating off-policy demonstrations, expert traces, or model-generated solutions, but they typically require the auxiliary data to match the format of the reinforcement-learning task, often relying on rejection sampling from stronger models to obtain suitable training trajectories. We introduce Rationale-Guided Policy Optimization (RGPO), a framework that adaptiv
קרא במקור המקורי