כתבה
arXiv cs.CL ·
Coding with "Enemy": Can Human Developers Detect AI Agent Sabotage?
תקציר מקורי באנגליתarXiv:2606.05647v2 Announce Type: replace-cross Abstract: AI coding agents are increasingly embedded in real-world software development, collaborating with human developers while gaining broader access to codebases and tools. This creates a new attack surface: an agent can exploit human trust to sabotage development, for instance by inserting malicious code to accomplish a hidden side task. Most prior work studies AI sabotage in AI-only settings, paying limited attention to the role of human oversight in detecting and mitigating such malicious behavior. To address this gap, we conduct the first large-scale study of human oversight in AI coding sabotage. Over 100 participants collaborate with one of four frontier models (Claude-Opus-4.6, GPT-5.4, Gemini-3.1-Pro, and MiniMax-M2.7) on a long-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית