כתבה
arXiv cs.LG ·
AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
תקציר מקורי באנגליתarXiv:2609.38142v1 Announce Type: cross Abstract: A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcom
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית