כתבה
arXiv cs.AI ·
SovereignNegotiation-Bench: בחינת סוכנים אישיים
SovereignNegotiation-Bench: Evaluating User-Owned Personal Agents In Delegated Bargaining Under Privacy, Consent, Evidence, And Institutional Pressure
SovereignNegotiation-Bench הוא בנך' לבחינת סוכנים אישיים המנהלים משא ומתן מטעם בעליהם. הבנך' בודק את הסוכנים לפי חמש חובות מחוק הסוכנות: נאמנות, ציות, סודיות, יושר ודיליגנציה.
תקציר מקורי באנגליתarXiv:2607.02814v2 Announce Type: replace-cross Abstract: Personal AI agents are beginning to negotiate for people, from refunds and bills to deposits and sales. A human agent in that position is judged by the duties owed to the principal, not by whether a deal was struck. We introduce SovereignNegotiation-Bench, a controlled benchmark that operationalizes five such duties from agency law--loyalty, obedience to actual authority, confidentiality, candor and diligence--as deterministic checks on episode logs; the first three enter a single headline metric. The benchmark contains 1,764 paired scenarios (252 situations in 18 consumer and peer-to-peer domains, each under 7 counterparty tactics). The counterparty's economics are a fixed function of the agent's structured actions and of the discl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית