כתבה
arXiv cs.LG ·
אמון בביקורת
Trust the Critic More
אלגוריתם RL חדש מסיר את הצורך להריץ כל נתיב לסיומו. האלגוריתם, AC2, מעביר קרדיט לקטעי פעולה, ולא לטוקן יחיד. זה נעשה על ידי שימוש בביקורת ראשונית, שמספקת דירוג עדכני של המצב. AC2 מציע פתרון חדש לבעיית קרדיט עבור LLMs, ומספק תוצאות טובות יותר באימון.
תקציר מקורי באנגליתarXiv:2609.39247v2 Announce Type: new Abstract: Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית