יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בדיקת עמידות של חוקרי שקרים: כשלי דימוי והתקשרויות חסרות משמעות

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations
בדיקת עמידות של חוקרי שקרים: כשלי דימוי והתקשרויות חסרות משמעות. נמצא כי חוקרי שקרים רבים נכשלים בזיהוי שקרים כאשר ה-LLM מקבל פרסונה אנטי-פקטית.
תקציר מקורי באנגליתarXiv:2609.39807v1 Announce Type: new Abstract: Lie detection probes aim to predict from a language model's internal states whether its output is truthful or dishonest. However, role-play complicates what "truth" means for an LLM: language models can adopt a wide range of personas that take very different claims to be true, including personas whose beliefs clearly contradict reality, such as a conspiracy theorist. In this work, we investigate whether lie detection probes reliably flag falsehoods generated under such an anti-factual persona or whether they instead follow the persona's beliefs. We introduce a dataset of 8,916 human-reviewed, on-policy responses from three LLMs adopting anti-factual personas. Evaluating eight probes from prior work, we find that many fail in this setting, par
קרא במקור המקורי