כתבה
arXiv cs.AI ·
Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability
תקציר מקורי באנגליתarXiv:2609.38712v1 Announce Type: new Abstract: Long-horizon agentic workflows require models to sustain repeated state-dependent actions all while the context grows, sub-task complexity changes, and new data arrives. Each situation represents an independent axis along which an agent may fail. An agent reconciling a long ledger, for example, must repeatedly read its state, update the correct record, and preserve alignment across thousands of outputs. A model may accept the entire ledger yet lose its place or stop applying the operation consistently as generation proceeds. We introduce Long-Transduction, a controlled diagnostic that tests a model's ability to stay on task during long generation while continuously reading, mutating, and outputting input-context dependent operations such as a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית