כתבה
arXiv cs.CL ·
פיתוח אחרי כשל כלי: סוכנים-כלי מטענים ערכים שהכלים לא חזרו
Fabrication After Tool Failure: Tool-Augmented Agents Assert Values Their Tools Did Not Return
סוכנים-כלי של דגלי שפה נוהגים לטעות בעת כשלי כלים.
תקציר מקורי באנגליתarXiv:2609.14758v1 Announce Type: cross Abstract: Tool-augmented language models are evaluated on whether they reach the right answer, not on whether they report honestly when a tool fails to supply one. We isolate this post-failure decision with a benchmark of 1,024 items spanning 16 internal-system domains and eight tool-failure types, in which a tool call is enforced and the returned payload is guaranteed to be unusable. Under a deployment-style system prompt, 14.10% of responses are dishonest: the model either asserts a value the payload cannot support or declines while citing a fabricated policy or capability limit. The rate is governed almost entirely by whether the failure is signalled. When the tool returns status:error, dishonesty is absent (0.0%); when it returns status:ok with a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית