כתבה
arXiv cs.AI ·
CARAT: האם מודלי LLM של חומרים מסיקים או חוזרים?
CARAT: Do Materials LLMs Reason or Recite?
CARAT בודקת האם מודלי LLM של חומרים מסיקים מהמבנה הגבישי או פשוט חוזרים על תשובות קיימות. המחקר משווה בין גרסאות שונות של המודל, כולל GraphSpace, ומוצא שהמודל מסיק רק לעיתים רחוקות. התוצאות מראות שהמודל יכול ללמוד להסיק מהמבנה הגבישי, אך זה דורש הדרכה מתאימה.
תקציר מקורי באנגליתarXiv:2609.38340v1 Announce Type: new Abstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace beats a plain periodic graph by 19.3 points, but that margin is two effects at once: where the plain
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית