כתבה
arXiv cs.CL ·
Tasks over Application Manuals: Revealing Gaps in Long-Horizon Procedural Reasoning for Language Models
תקציר מקורי באנגליתarXiv:2609.13005v1 Announce Type: new Abstract: Large language models (LLMs) have achieved strong performance on a wide range of natural language tasks, and recent benchmarks suggest that they are increasingly adept at multi-hop reasoning. However, these benchmarks are typically short-horizon, requiring only a small number of retrieval or inference steps, and provide limited evidence of reliability on real-world tasks that involve following manuals spanning hundreds of pages with complex, interdependent guidelines. In this paper, we introduce Tasks over Application Manuals (TAM), a benchmark for evaluating long-horizon procedural reasoning. We construct TAM by curating real-world tasks from two domains: ICD-10-CM clinical coding (mapping medical conditions to diagnostic codes) and U.S. fed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית