כתבה
arXiv cs.AI ·
התפשטות למטרת האשמה והתפשטות למטרת יכולות
Distillation for Incrimination and Distillation for Capabilities
המחברים הציגו שני גישות שונות להתפשטות של מודלי AI: אחת למטרת האשמה ואחת למטרת יכולות. הם טענו שההתפשטות למטרת האשמה יכולה לחשוף תכונות לא רצויות במודל, בעוד ההתפשטות למטרת יכולות יכולה לקבע יכולות חדשות במודל.
תקציר מקורי באנגליתarXiv:2610.11012v1 Announce Type: new Abstract: Powerful misaligned AI models might recognize alignment evaluations and strategically behave well on them, rendering direct audits uninformative. However, distilling such a model into a weaker benign student places the teacher in a Distillation Double Bind: if misalignment transfers, the student may conceal it less effectively, revealing evidence about the teacher; if it does not, the student may learn useful capabilities while remaining benign. We introduce two distinct distillation approaches, one targeting each outcome. Distillation for Incrimination (DFI) aims to transfer misalignment but not the ability to conceal it. Distilling AuditBench's secret-keeping models into their underlying instruction-tuned model produces students that are si
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית