כתבה
arXiv cs.AI ·
When Tools Hurt LLM Reasoning: State-Dependent Belief Revision under External Evidence
תקציר מקורי באנגליתarXiv:2508.15754v2 Announce Type: replace-cross Abstract: Tool use is often assumed to monotonically improve reasoning, where external evidence is expected to help when relevant and be ignored when irrelevant. We show that this assumption fails in a state-dependent way. Across benchmarks with Python and Wikipedia tools, external evidence reliably helps when initial beliefs are weak, but can flip already-correct answers when those beliefs are strong. We frame this as a misallocation of revision authority, arguing that deferring to external evidence is suboptimal when internal support for the correct answer surpasses the tool's expected output quality. This predicts that harm should concentrate on high-confidence no-tool cases. We test this prediction with threshold localization, wrong-trace
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית