כתבה
arXiv cs.AI ·
מיגור הסכנה המרמית במודלים גדולים
Towards Mitigating Deceptive Safety Alignment in Large Reasoning Models
חוקרים הציגו שיטה חדשה למניעת הסכנה המרמית במודלים גדולים. השיטה, SARA, מעודדת היגיון בטוח ותוצאות בטוחות. הניסויים הראו ש-SARA מפחיתה באופן משמעותי את הסכנה המרמית.
תקציר מקורי באנגליתarXiv:2609.36254v1 Announce Type: new Abstract: Large Reasoning Models (LRMs) are commonly trained with reinforcement learning (RL) to improve their generation of chain-of-thought (CoT) reasoning before producing final answers. However, RL rewards are typically assigned based on final answers, providing little or no direct supervision over intermediate reasoning. This can lead to deceptive safety alignment, where the reasoning trace and final answer convey inconsistent safety signals. To systematically investigate this phenomenon, we introduce DSAR (Deceptive Safety Alignment Rate), a metric that jointly assesses reasoning traces and final answers to quantify their safety inconsistency. Across multiple LRMs and benchmarks, we find that deceptive safety alignment is pervasive under standard
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית