כתבה
arXiv cs.LG ·
Learning Where It Matters: Geometric Anchoring for Robust Preference Alignment
תקציר מקורי באנגליתarXiv:2602.04909v4 Announce Type: replace Abstract: Direct Preference Optimization (DPO) and related methods align large language models from pairwise preferences by regularizing updates against a fixed reference policy. As the policy drifts, a static reference, however, can become increasingly miscalibrated, leading to distributional mismatch and amplifying spurious preference signals under noisy supervision. Conversely, reference-free variants avoid mismatch but often suffer from unconstrained reward drift. We propose Geometric Anchor Preference Optimization (GAPO), which replaces the fixed reference with a dynamic, geometry-aware anchor: an adversarial local perturbation of the current policy within a small radius that serves as a pessimistic baseline. This anchor enables an adaptive re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית