יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הבובנאי הנסתר: חיזוי שינוי אמונות אנושיות בדיאלוגים מניפולטיביים של LLM

The Hidden Puppet Master: Predicting Human Belief Change in Manipulative LLM Dialogues
חוקרים פיתחו מסגרת תאורטית לחיזוי שינוי אמונות אנושיות בדיאלוגים מניפולטיביים של LLM. המחקר מראה כי מודלים יכולים לחזות שינויי אמונות, אך עם הטיות משמעותיות. החידוש יכול לשמש לשיפור בטיחות AI.
תקציר מקורי באנגליתarXiv:2603.20907v4 Announce Type: replace Abstract: As users increasingly turn to LLMs for practical and personal advice, they become vulnerable to subtle steering toward hidden incentives misaligned with their own interests. While existing NLP research has benchmarked manipulation detection, these efforts often rely on simulated debates and remain fundamentally decoupled from actual human belief shifts in real-world scenarios. We introduce PUPPET, a theoretical taxonomy and resource that bridges this gap by focusing on the moral direction of hidden incentives in everyday, advice-giving contexts. We provide an evaluation dataset of N=1,035 human-LLM interactions, where we measure users' belief shifts. Our analysis reveals a critical disconnect in current safety paradigms: while models can
קרא במקור המקורי