כתבה
arXiv cs.LG ·
אופטימיזציה רציפה של תועלת ישירה
Continuous-Utility Direct Preference Optimization
פרקטיקה חדשה להסברה על-מודלי שפה גדולה, המשפרת את תפיסת ההסברה המדויקת.
תקציר מקורי באנגליתarXiv:2602.00931v3 Announce Type: replace Abstract: Large language model reasoning is often treated as a monolithic capability, relying on binary preference supervision that fails to capture partial progress or fine-grained reasoning quality. We introduce continuous utility direct preference optimization (CU-DPO), a framework that aligns models to a portfolio of prompt-based cognitive strategies by replacing binary labels with continuous scores that capture fine-grained reasoning quality. We prove that learning with K strategies yields a Theta(K log K) improvement in sample complexity over binary preferences and that DPO converges to the entropy-regularized utility-maximizing policy. To exploit this signal, we propose a two-stage pipeline: (i) strategy selection, which optimizes the model
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית