יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

תן אמון במבקר

Trust the Critic More
AC2 משפר את תהליך האימון של מודלי שפה גדולים. הוא מאפשר הקצאת אשראי לפעולות קצרות, במקום להמתין לתוצאה סופית. המחברים השתמשו ב-AC2 כדי לאמן את Qwen3-4B על FineProofs-RL.
תקציר מקורי באנגליתarXiv:2609.39247v1 Announce Type: cross Abstract: Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods can provide finer-grained credit assignment, but learned critics are generally considered too inaccurate to trust when training LLMs with RL. In recent works, even when a critic is present, it is used only for baseline estimation, so every trajectory must be rolled out to its terminal reward. We introduce Actor-Critic with Action Chunking (AC2) that removes the need to roll every trajectory to completion. AC2 instead assigns credit to action chunks: short continuations of prefixes of past trajectories. A learned critic scores the state reached at the end of each action chunk, allowing the
קרא במקור המקורי