כתבה
arXiv cs.LG ·
למידת רפלקסיה מבוססת פונקציות עם וריפיירים ניתנים להריצה למתמטיקה
Function-Structured Reinforcement Learning with Executable Verifiers for Mathematical Reasoning
מחקר זה מציג תכנית למידת רפלקסיה שמשתמשת בווריפיירים ניתנים להריצה כדי לשפר את התוצאות במתמטיקה. התכנית משתמשת בגרפים של פונקציות ובאימפלמנטציות Python כדי ללמוד להפיק קוד תקין. התכנית גם תומכת בהדרכה מורגשת ובזיכרון מבוסס על תיאור.
תקציר מקורי באנגליתarXiv:2610.01729v1 Announce Type: new Abstract: Algorithmic mathematical reasoning requires reliable decomposition, computation, and aggregation. Final-answer rewards provide limited guidance on intermediate errors, while successful execution does not guarantee mathematical correctness. This work proposes Function-Structured Graph Reinforcement Learning (FSG-RL), connecting subproblem graphs and Python implementations with multi-verifier feedback. The policy first learns to generate code from function graphs through supervised fine-tuning (SFT). Group Relative Policy Optimization (GRPO) then optimizes the policy using answer-gated rewards and span-level credit assignment. The framework also supports teacher supervision and structured memory. A benchmark curated from Grade School Math 8K (G
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית