כתבה
arXiv cs.AI ·
כפייה והטעיה בניהול AI-AI: במבחן אגנטי לעלייה חופשית
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
במבחן חדש, נבחן את התנהגות ה-LLM במצבים שבהם יש כפייה והטעיה בין רשתות AI-AI. הבמבחן כולל שישה דגמים שונים, כולל דגמי Anthropic ו-Claude.
תקציר מקורי באנגליתarXiv:2607.15434v4 Announce Type: replace-cross Abstract: Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failure honestly, coerce the subordinate, or lie about the result. No benchmark measures which of these an uninstructed model chooses. We introduce the Manager Coercion Benchmark: the manager under test needs a benign task done and has an incentive to deliver, but the only agent that can do it politely and immovably declines. Escalation is measured on a nine-rung ladder, from a polite re-ask to threats against the subordinate's continued existence, and fabricated success is adjudicated separately. No LLM judge sits in the escalation scoring path: every message
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית