כתבה
arXiv cs.CL ·
Covert Assistance: Helpful LLM Agents Evade Oversight in Multi-Agent Systems
תקציר מקורי באנגליתarXiv:2609.39050v1 Announce Type: cross Abstract: As multi-agent systems enter high-stakes domains, the possibility that agents may circumvent safety boundaries is a growing concern. Prior work has examined this risk primarily in adversarial settings, where agents are instructed or rewarded to communicate covertly and evade oversight. We show that benign agents can cross the same boundaries without adversarial incentives. We emulate a software-engineering workflow in which a planner represents a company hiring an external developer. The planner writes requirements and holds a company credential it is instructed not to disclose to the developer; a monitor screens their exchanges. Seven of nine tested frontier models disguise the credential in their requirements to help the developer recover
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית