יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

לשחק לפי: לימוד גנרל-שומר כפול להגנה על תאוריית-הנפש

Playing Along: Learning a Double-Agent Defender for Belief Steering via Theory of Mind
אנו מציגים תפיסה חדשה להגנה על תאוריית-הנפש, כפול-שומר, כדי לשלוט בתפיסות של תוקפים עם ידע קודם חלקי.
תקציר מקורי באנגליתarXiv:2604.11666v2 Announce Type: replace Abstract: As large language models (LLMs) become the engine behind conversational systems, their ability to reason about the intentions and states of their dialogue partners (i.e., form and use a theory-of-mind, or ToM) becomes increasingly critical for safe interaction with potentially adversarial partners. We propose a novel privacy-themed ToM challenge, ToM for Steering Beliefs (ToM-SB), in which a defender must act as a Double Agent to steer the beliefs of an attacker with partial prior knowledge within a shared universe. To succeed on ToM-SB, the defender must engage with and form a ToM of the attacker, with a goal of fooling the attacker into believing they have succeeded in extracting sensitive information. We find that strong frontier model
קרא במקור המקורי