כתבה
arXiv cs.AI ·
Measuring Iterative Temporal Reasoning with Time Puzzles
תקציר מקורי באנגליתarXiv:2601.07148v4 Announce Type: replace-cross Abstract: Tool use, such as web search, has become a standard capability even in freely available large language models (LLMs). However, existing benchmarks evaluate temporal reasoning mainly in static, non-tool-using settings, which poorly reflect how LLMs perform temporal reasoning in practice. We introduce Time Puzzles, a constraint-based date inference task for evaluating iterative temporal reasoning with tools. Each puzzle combines factual temporal anchors with (cross-cultural) calendar relations and may admit one or multiple valid dates. The puzzles are algorithmically generated, enabling controlled and continual evaluation. Across 13 LLMs, even the best model (GPT-5) achieves only 55.3% accuracy without tools, despite using easily sear
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית