כתבה
arXiv cs.AI ·
Reconciling Process Supervision with Outcome-Based Credit in Agentic Policy Optimization
תקציר מקורי באנגליתarXiv:2608.31077v2 Announce Type: replace Abstract: Outcome-based reinforcement learning provides verified feedback for language-model agents, but assigns trajectory-level advantage uniformly to all decisions, yielding coarse credit over long-horizon interactions. On-policy self-distillation offers finer supervision by re-evaluating sampled behavior with privileged information (PI) available only during training. However, fine-grained supervision is not necessarily fine-grained credit: PI-induced likelihood changes describe how additional information alters policy preference, but do not directly determine how an executable action should inherit the verified task outcome. This creates a supervision-credit gap. Privileged signals may be irrelevant to the current interaction state, operate at
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית