יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CM-DPO: Constraint-Margin Direct Preference Optimization for LLM Planning

תקציר מקורי באנגליתarXiv:2610.09219v1 Announce Type: cross Abstract: Direct Preference Optimization (DPO) treats all constraint violations equally: a $1 budget overshoot and a $1,000 overshoot induce the same training signal. It is also susceptible to length and style bias when preference pairs come from different model families. We introduce Constraint-Margin DPO (CM-DPO), which replaces DPO's binary preference signal with a continuous margin derived from a deterministic symbolic verifier and scaled by violation severity. Hard and soft constraints are separated through a lexicographic objective, ensuring hard constraints are never traded off against preferences. To supply CM-DPO with bias-reduced training pairs, we generate preference data through procedurally generated constraint profiles (DCCG) and minima
קרא במקור המקורי