כתבה
arXiv cs.CL ·
כיוונון העדיפויות כארגון מחדש של עדכון ספקטרלי
Preference Tuning as Spectral Update Reorganization
חוקרים את RLHF ואופטימיזציה של העדיפויות דרך מבנה ספקטרלי. התוצאות מראות כי עדכונים מתארגנים בצורה מובנית, עם 'ראש' קומפקטי האחראי לשינוי ההתנהגותי. ה'זנב' הוא הטרוגני ותורם מעט להתנהגות.
תקציר מקורי באנגליתarXiv:2607.20438v1 Announce Type: new Abstract: Preference-based post-training is usually understood through endpoint behavior, yet the learned update that produces this behavior remains largely opaque. We study RLHF and related preference optimization through the spectral structure of their induced parameter updates. By decomposing effective LoRA updates and reloading their spectral components as plug-in modules, we turn preference-induced updates into objects that can be isolated, recomposed, and directly intervened on. Across model families, optimization algorithms, and supervision regimes, these updates consistently develop a spectral head--tail organization. A compact head emerges early and carries the dominant endpoint shift, while a heterogeneous residual tail remains. The split is
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית