כתבה
arXiv cs.AI ·
SciTrek: Evaluating and Improving Long-Context Numerical Reasoning over Scientific Articles
תקציר מקורי באנגליתarXiv:2509.21028v5 Announce Type: replace Abstract: We introduce SciTrek, a synthetic question-answering dataset for assessing and improving long-context numerical reasoning in large language models (LLMs). Existing long-context datasets with inputs beyond 64K tokens either target simple information retrieval or, when they do involve reasoning, rely on artificial contexts, while numerical reasoning remains largely overlooked in both cases. SciTrek addresses these limitations with questions that require numerical operations (e.g., counting, sorting, aggregation, and comparison) over collections of full-text scientific articles. Questions are generated automatically by formulating them as SQL queries over a database of article metadata (titles, authors, and references), and ground-truth answ
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית