יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מידע כוזב ללא טריגרים

Misinformation Without Triggers: From Factual Answers to Downstream Decisions
מודלי שפה יכולים להפיק תשובות כוזבות, אפילו כאשר התשובה הישירה נכונה. מחקר זה בודק את הפער בין תשובות ישירות לבין החלטות שנובעות מהן, ומוצא כי תיקון עובדות אינו תמיד מספיק כדי למנוע החלטות שגויות.
תקציר מקורי באנגליתarXiv:2610.02886v1 Announce Type: cross Abstract: Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a m
קרא במקור המקורי