יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לאחריות: שיפור פוליטי קונסרבטיבי עם שיטת Cross-Entropy ל-RFT

Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
שיטה חדשה לשיפור רכיבי רפורמציה חופשיים למודלי שפה גדולים, המשתמשת בשיטת Cross-Entropy. זוהי שיטה חופשית מביקורת, המאפשרת שיפור רכיבי רפורמציה חופשיים למודלי שפה גדולים. השיטה נבחנה במספר תחומים, כולל סוקובאן וחיפוש-R1.
תקציר מקורי באנגליתarXiv:2610.03361v1 Announce Type: new Abstract: Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textit{Follow the Winners} (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW through
קרא במקור המקורי