כתבה
arXiv cs.CL ·
ללא סימן: מענייניות פושטת רגל ללא סימן
Misinformation Without Triggers: From Factual Answers to Downstream Decisions
דגימות שגויות יכולות להפיץ מידע שגוי בלי סימן. נמצא פער בין התשובה הישירה לבין ההחלטה שהיא יוצרת. נמצא כי תשובה ישירה לא תמיד מבטיחה החלטה ישירה.
תקציר מקורי באנגליתarXiv:2610.02886v1 Announce Type: new Abstract: Language models learn from web documents, some of them false, and false content can reach a model's answer to a factual question and the summaries and decisions that use it. Most data-poisoning studies add a trigger to the training data and activate it in the prompt. False documents can also change factual responses without any trigger, but we do not know whether the direct answer predicts the decision. In this work, we follow false content past the answer and find an \emph{audit gap} between what a direct probe reports and what the model then does, comparing false training with matched truthful controls in a controlled decision task, \emph{Guess the Capital}, where a fixed decoder turns factual answers into a scored card choice, and on a mis
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית