כתבה
arXiv cs.AI ·
שחזור ידע מדעי מרומז
Reconstructing Implicit Scientific Knowledge: Evaluating LLM Agents through End-to-End Reproduction of Astronomy
חוקרים בדקו את יכולתם של מודלי LLM לשחזר ידע מדעי מרומז. הם פיתחו שיטה לבדיקת סוכנים באמצעות שחזור מלא של מחקרים אסטרונומיים.
תקציר מקורי באנגליתarXiv:2609.35900v1 Announce Type: cross Abstract: The integration of large language models (LLMs) into scientific workflows is accelerating, yet their ability to reconstruct the reasoning underlying published research remains unexplored. Papers specify explicit procedures while leaving many methodological dependencies-data selection, calibration corrections, priors, and domain assumptions-implicit. This ambiguity complicates the evaluation of LLM-based agents, since a failure to reproduce a result may reflect either limitations of the agent or underspecification in the source. We present a framework that evaluates agents through end-to-end reproduction, separating execution from verification and computational failure from methodological ambiguity. We apply it to fourteen astronomy studies:
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית