כתבה
arXiv cs.AI ·
בדיקת עמידות של LLM לזיהוי שקרים: כשלי דימוי ותלותיות
Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations
במאמר זה, נחקרה עמידות של LLM לזיהוי שקרים. נמצא כי רבים מהמבחנים כושלים בסצנת דימוי, וכי רבים מהם עוקבים אחר רעיונות שהם ספורטיביים באופן שקרי. נציגה חדשה של מבחן רגיל, שהיא פשוטה ומבוססת על קווי עקבה, הציגה את הביצועים הטובים ביותר בשני הבדיקות.
תקציר מקורי באנגליתarXiv:2609.39807v1 Announce Type: cross Abstract: Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית