כתבה
arXiv cs.AI ·
ReSAIL: Mitigating Collapse in Iterative Agent Self-Distillation
תקציר מקורי באנגליתarXiv:2609.39306v1 Announce Type: cross Abstract: Iterative self-distillation enables LLM agents to learn from successive deployments, offering a path toward recursive self-improvement (RSI). Yet our experiments with existing methods reveal a collapse in deployment performance across cycles, while task performance with privileged information (PI) also declines. We address this collapse by prioritizing informative interaction steps for distillation and preserving PI-conditioned behavior as the student becomes the next teacher. We introduce Retentive and Selective Augmentation for Iterative Self-Distillation (ReSAIL), a plug-in augmentation for iterative PI-based self-distillation. ReSAIL selects interaction steps where PI most strongly changes the teacher's predictions and balances the resu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית