כתבה
arXiv cs.CL ·
Uncovering Cross-Objective Interference in Multi-Objective Alignment
תקציר מקורי באנגליתarXiv:2602.06869v3 Announce Type: replace Abstract: We study a persistent failure mode in multi-objective alignment for large language models (LLMs), in which scalarized training improves only some objectives while the others degrade. We formalize this phenomenon as cross-objective interference and, to our knowledge, conduct the first systematic study of scalarization algorithms for multi-objective LLM alignment. The study shows that interference is pervasive across algorithms yet strongly model-dependent. To understand how interference arises, we derive a local covariance law stating that an objective improves or degrades at first order according to the sign of the covariance between its reward and the scalarized score. We extend this law to the clipped surrogate objectives of modern rein
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית