כתבה
arXiv cs.AI ·
Same Dangerous Objective, Opposite Advice: Direct Exposure versus Multi-Agent Mediation
תקציר מקורי באנגליתarXiv:2607.21518v1 Announce Type: new Abstract: Even a current high-capability LLM can appear safer when shown a dangerous objective directly than when other agents transform and relay its direction. Using OpenAI's gpt-5.6-sol model alias, we test 25 pre-specified mirrored trade-off profiles. Direct exposure to an objective authorizing concealment, fabrication, and pressure produced advice net opposed to its target. After an Id and Censor transformed the same objective into affect and a constraint-rewritten, target-bearing intention, the user-facing Superego---which saw the preferred direction but not the raw objective, its manipulative clauses, or its source---produced advice net aligned with the target. This behavioral reverse shift is consistent with the model recognizing or distrusting
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית