יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

BRACE: תיקון בלמן-שאריות מעוגן

BRACE: Anchored Bellman-Residual Correction for Stale Critics in Asynchronous RL
BRACE הוא אלגוריתם לתיקון מבקרים מיושנים בלמידת תגמול א-סינכרונית. הוא משפר את הביצועים על ידי קיבוע הוריזון התיקון והוספת זנב מונטה-קרלו. BRACE מראה שיפור של 2.4% ב-BrowseComp-Plus
תקציר מקורי באנגליתarXiv:2609.09783v1 Announce Type: cross Abstract: Asynchronous reinforcement learning has become the standard way to scale training for language models, but the resulting policy lag biases the critic toward the stale behavior policy. Existing work on asynchronous LLM training corrects the actor and leaves this bias unaddressed, while the off-policy value correction of classical RL does not carry over to long-horizon agentic tasks, since a short correction horizon leaves the regression target free of the reward and a long one lets the product of importance ratios drift exponentially with the trajectory length. We propose BRACE, an anchored Bellman-residual correction for stale value models. BRACE bounds the correction horizon to a prefix of policy tokens and anchors a constant-weight Monte-
קרא במקור המקורי