כתבה
arXiv cs.AI ·
LOCUS: תאימות-משימתית לדרגה נמוכה לאחר-אימון לשפע תגובות-טוקן
LOCUS: Task-Aware Low-Rank Post-Training for Token-Efficient Language Generation
LOCUS היא שיטה שבוחרת תת-מרחב תאימות-משימתי נמוך-דרגה לאחר-אימון שמקטין את אורך התגובה ללא הפסד באיכות. השיטה נבדקה על ידי Anthropic.
תקציר מקורי באנגליתarXiv:2609.11739v1 Announce Type: cross Abstract: Large language model serving costs scale directly with output sequence length, yet standard preference alignment often inflates response verbosity without improving utility. We study whether the parameterization of post-training updates affects generation length: low-rank subspaces alter sequence length without modifying the alignment loss. We present LOCUS, a method that selects a task-aware low-rank adaptation subspace to minimize output-token cost subject to a utility constraint. Within this subspace, post-training retains the native preference objective with a frozen backbone. Across Anthropic HH-RLHF dialogue preferences, we evaluate two $\sim$3B decoder backbones, Pythia-2.8B and Qwen2.5-3B, against protocol-matched full-parameter DPO
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית