יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיפור מדיניות שמרני עם שיטת ה-Entropy ל-RFT

Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
חוקרים מציעים אלגוריתם חדש לשיפור מדיניות במודלים גדולים של שפה, באמצעות שיטת ה-Entropy. האלגוריתם, Follow the Winners, מאפשר לבצע שיפורים במדיניות ללא צורך במודל ביקורת. המחקר מראה כי האלגוריתם מתאים לסביבות מציאותיות ויכול להחליף את השיטות הקיימות.
תקציר מקורי באנגליתarXiv:2610.03361v1 Announce Type: cross Abstract: Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services or security sandboxes, where repeated rollouts are impractical to obtain and aggressive updates entrench the noise of long, sparsely verified trajectories. We propose \textit{Follow the Winners} (FTW), a critic-free policy-learning algorithm that adapts the cross-entropy method to RFT, replacing group rollouts with an ordinal filter on replay-buffer samples that yields polynomial concentration in the order statistic of returns. We derive FTW throu
קרא במקור המקורי