כתבה
arXiv cs.AI ·
Lightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition
תקציר מקורי באנגליתarXiv:2609.14708v1 Announce Type: new Abstract: A core goal of efficient reasoning is to improve the accuracy-efficiency frontier. However, jointly improving reasoning accuracy and inference efficiency can be challenging, as the two objectives can favor different reasoning behaviors. Independently post-trained models already offer distinct strengths in accuracy and efficiency. We introduce Lightning Weave, a post-training framework that extracts and composes these independently learned capabilities in a single student through on-policy distillation. Each acquired capability is represented by the policy shift from the model before post-training to the resulting specialist. Lightning Weave combines aligned log-ratio shifts at shared student token states and uses Tilted-Target DOPD to convert
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית