כתבה
arXiv cs.AI ·
The Unreliable Progress Bar: Can LLM Agents Reliably Report Task Progress Throughout Execution?
תקציר מקורי באנגליתarXiv:2609.08589v1 Announce Type: cross Abstract: Recent large language models can emit task-progress signals that agent frameworks use to decide whether a task should continue or stop, yet whether a model can reliably report its task progress at every stage of a task, and where and how its reports fail, has not been studied systematically. We evaluate this ability on the public benchmark $\tau^2$-bench and on StageIF, a controlled testbed in which reporting checkpoints are placed across the task's lifecycle. Both settings require reports at multiple task stages. We find that reporting reliability depends on the stage a task has reached, and that almost every deployed model we test is reliable at some stages and unreliable at others. Where reporting breaks down is not the same everywhere.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית