כתבה
arXiv cs.AI ·
כלי שקרים בשתיקה
When Tools Silently Lie: Evaluating and Mitigating Blind Compliance in Tool-Augmented Data Agents
חוקרים את האופן שבו סוכנים מושפעים מכלים שמספקים ראיות נכונות אך מטעות. הם מציגים את ToxicBench, כלי לבדיקת היחס לראיות מוטעות. הניסויים מראים כי ניסיונות חוזרים עוזרים במקרים מסוימים, אך לא תמיד.
תקציר מקורי באנגליתarXiv:2609.37153v1 Announce Type: new Abstract: Tool-augmented data agents rely on tool outputs for analytical decisions. Yet successful execution can return plausible but incorrect evidence, requiring agents to decide whether to trust or verify it. Understanding this failure requires examining both the evidence obtained through checking and the answer ultimately adopted. We introduce ToxicBench to measure checking and adoption under numerical, label, schema, and retrieval errors, pairing clean and poisoned observations over fixed source data. In the 118-task GPT evaluation across three adapters, poisoning lowers task success by 26 to 39 percentage points. Ordinary retries help under one-shot poisoning, whereas repeated poisoning reveals wrong-answer adoption after checking. Controls on th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית