כתבה
arXiv cs.LG ·
TD-DPO: אופטימיזציה של העדפות תלויות הבדלים
TD-DPO: Difference-Aware Preference Optimization for Mitigating Sycophancy in Clinical Autism Intervention Dialogue
TD-DPO היא שיטה לאופטימיזציה של העדפות תלויות הבדלים, שנועדה להפחית את הסיכון הבטיחותי בדיאלוגים טיפוליים לילדים אוטיסטים. השיטה משתמשת באופטימיזציה של העדפות ברמת הטוקן, כדי לשפר את היכולת של מודלים להבין ולתגוב להעדפות המשתמש.
תקציר מקורי באנגליתarXiv:2607.18304v1 Announce Type: new Abstract: The sycophancy of large language models can increase the safety risk in intervention dialogue for autistic children. Supervised fine-tuning can somewhat reduce sycophancy, but relying solely on positive examples is often insufficient to identify and correct failure patterns. We observe that sycophancy behaviors can often be localized to a limited span within the model response. In this regime, sequence-level preference optimization can over-update preference-irrelevant tokens and degrade intervention ability. To address this, we propose the \textbf{M}inimal \textbf{E}dit \textbf{D}ata \textbf{A}ugmentation (MEDA) strategy to construct controlled, stable, minimal edit preference pairs and \textbf{T}oken-level \textbf{D}ifference \textbf{D}irec
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית