כתבה
arXiv cs.LG ·
PR-OPD: Privileged Representation On-policy Self-Distillation for Agentic Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.36642v1 Announce Type: new Abstract: Language-model agents are usually trained by reinforcement learning from one reward per episode, and privileged self-distillation enriches it by letting the same policy, given a skill, teach its skill-free self through token probabilities. However, we identify two phenomena that question this channel. Invisible Advantage: a skill in context lifts WebShop success from 42.2% to 56.2%, yet changes the probabilities of fewer than a quarter of the sampled tokens. Much to Align: a skill changes the hidden states of over 80% of response tokens, in a way that linear probes can trace back to the specific skill. To exploit this, we propose Privileged Representation On-policy Self-Distillation (PR-OPD). After a GRPO warm start, the policy writes a hinds
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית