כתבה
arXiv cs.AI ·
בניית בסיסי מבחנים לבדיקת תקינות הסיבות של LLM ב-AI מדעי
Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
במאמר זה, צוות מדענים פיתחו בסיסי מבחנים חדשים לבדיקת תקינות הסיבות של LLM ב-AI מדעי. הבסיסי כולל 33,819 מבחנים שנבנו באופן אוטומטי מאונטולוגיות OWL 2. המבחנים נועדו לבדוק את יכולת ה-LLM לסבול על ידי פרטי ידע, ולהציג תקינות סיבתית. הבסיסי כולל גם 6 LLMs שהוכנסו זריז, והציגו תוצאות של 41.1-76.8% נכונות. המחברים טוענים שהבסיסי יכול לשמש ככלי לבדיקת תקינות הסיבות של LLM ב-AI מדעי.
תקציר מקורי באנגליתarXiv:2610.00682v1 Announce Type: new Abstract: Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiomat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית