יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

דג'ה וו מולקולרי: שחזור ערכים מפורסמים במודלים שפה חדישים

Molecular D\'ej\`a Vu: Digit-Level Retrieval of Published Values in Frontier Language Models
בדיקת 22 מודלים שפה גדולים על בסיס מחקר מולקולרי. נמצא כי המודלים משחזרים ערכים מפורסמים במקום לחזות אותם. הניסויים נערכו בשני רמות חשיבה והראו כי החשיבה משפיעה על השחזור.
תקציר מקורי באנגליתarXiv:2609.05381v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly evaluated on molecular property benchmarks, but accuracy cannot distinguish a model that predicts a property from one that retrieves a published number. We audit 22 frontier models on 12 regression benchmarks for verbatim retrieval and find that it is widespread but relatively benchmark-specific: on five datasets more than $50\%$ of the LLMs show verbatim retrieval, while on the remaining datasets it appears only in isolated cells. We run our experiments at two reasoning levels and find that reasoning changes retrieval. The same experiments, on the same molecules and with the same prompt, are flagged $89\%$ more often at the higher reasoning level than at the lowest one. Finally, we test a way to
קרא במקור המקורי