יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בנצ'מרקים לבדיקת LLM

Ontology-Grounded, Reasoner-Verified Benchmarks for Evaluating LLM Reasoning in Scientific AI
פותחו בנצ'מרקים חדשים לבדיקת מודלי LLM. הבנצ'מרקים מבוססים על אונטולוגיות ובודקים את היכולת של המודלים להגיע למסקנות נכונות. הבנצ'מרקים נועדו לשפר את הביצועים של מודלי LLM ביישומים מדעיים.
תקציר מקורי באנגליתarXiv:2610.00682v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly underpin scientific AI applications that reason over structured knowledge, from biomedical question answering to materials informatics. However, their logical reasoning often falls short, producing factual inaccuracies unacceptable in these settings. Reliable evaluation remains challenging: manual dataset construction scales poorly, and LLM-based generation risks embedding the very flaws it aims to measure. High-quality benchmarks must ground both correct and incorrect labelled examples in explicit background knowledge, formally verifiable by a standard reasoner. We propose a pipeline that automatically generates ontology-grounded multiple-choice question (MCQ) benchmarks from any sufficiently axiom
קרא במקור המקורי