כתבה
arXiv cs.AI ·
Molecular D\'ej\`a Vu: Digit-Level Retrieval of Molecular Properties in Frontier Language Models
תקציר מקורי באנגליתarXiv:2609.05381v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly employed to predict molecular properties. However, prediction error alone cannot distinguish prediction from retrieval of published values. We audit 22 frontier models on 12 molecular regression datasets in a zero-shot setting, assessed against a molecule-blind reference derived from each dataset's labels. Significant retrieval is concentrated on 5 datasets, with isolated flagged LLMs elsewhere. Increasing the reasoning setting raises the number of flagged model--dataset combinations from 47 to 89 of 264. An in-context blinding experiment reduces retrieval but leaves a quarter of the combinations flagged. Blinding changes model rankings and increases errors. Because blinding also removes chemi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית