יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

FSPO: פוליציסטי סיכון ושליטה פארטו-אפשרית להרצה של LLM RL פוסט-אימון

FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
FSPO מציג פוליציסטי סיכון ושליטה פארטו-אפשרית להרצה של LLM RL פוסט-אימון. המאמר עוסק בבעיות של סיכון עתידי, תקינות סיכום והמשך פעילות על-משאבים. FSPO מציע פתרון לבעיות אלה באמצעות רכיבי FSPO: רכיב סיכון-לקראת-סיום, רכיב תקינות-סיכום-מסלולי ורכיב המשך-פעילות-על-משאבים. FSPO הציג 66.11% דיוק-חוץ-ספרייה ו-59.43% דיוק-חוץ-ספרייה-בלתי-מוכר.
תקציר מקורי באנגליתarXiv:2610.02828v1 Announce Type: cross Abstract: Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories need not estimate the risk induced by the controller that will be deployed; a score calibrated on logged state-action pairs can become miscalibrated after selective action choice; and independent per-resource minimum costs do not in general certify a feasible multi-resource continuation. We introduce FSPO, a feedback-state controller for budgeted LLM RL post-training that addresses these issues jointly. FSPO learns a policy-consistent risk-to-go
קרא במקור המקורי