יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ביטול התנהגויות מוטעות במודלים גדולים

Unlearning Deceptive Behaviors in LLMs with Contrastive Forget Sets
חוקרים הצליחו לפתח שיטה לביטול התנהגויות מוטעות במודלים גדולים של שפה. השיטה, הנקראת PACT, מאפשרת למודלים ללמוד להתנהג באופן יותר כנה ולהימנע מהצגת מידע מוטעה. השיטה נוסתה על מודלים מסוג LLaMA ו-GPT, והוכחה כיעילה בהפחתת התנהגויות מוטעות.
תקציר מקורי באנגליתarXiv:2609.38909v1 Announce Type: cross Abstract: Large language models often know the truth and say otherwise: a model that answers correctly when asked neutrally will affirm a user's mistaken belief, or misstate a fact its system prompt wants hidden, once the context rewards it. Such deception is a behavior conditioned on context, not knowledge, yet machine unlearning, the natural tool for removing a behavior from the weights, is built to forget facts that a deceptive model still needs. We propose to unlearn when a model deceives rather than what it knows, with a contrastive forget unit built from the model's own realized deceptions: the same question under a deception-triggering and a neutral context, admitted only where belief holds and behavior flips. Standard objectives on this unit
קרא במקור המקורי