יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התגברות על התנהגויות שקריות ב-LLMs באמצעות קבוצות זיכרון נגדיות

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
במאמר זה, המחברים מציגים פתרון לבעיית ההתנהגויות השקריות של LLMs, על ידי שימוש בקבוצות זיכרון נגדיות. הם מציעים פתרון שמטרתו להתגבר על ההתנהגויות השקריות, כולל שמות מודלים/כלים/חברות.
תקציר מקורי באנגליתarXiv:2609.38909v1 Announce Type: new Abstract: Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit fa
קרא במקור המקורי