כתבה
arXiv cs.LG ·
CausalBN-Bench: תקן ניסוי ליכולת הלמידה הסיבתית של LLMs
CausalBN-Bench: A Comprehensive Benchmark for Causal Learning Capability of LLMs
CausalBN-Bench הוא תקן חדש לבדיקת יכולות הלמידה הסיבתית של LLMs. התקן כולל שלושה תרגילים שונים, כולל קורלציה, גוף סיבתי והזיהוי של סיבות. התקן נועד לבדוק את יכולות ה-LLMs להבין סיבות ולהסביר תוצאות. ה-CausalBN-Bench כולל גם ידע רקע ונתוני הכשרה בפני ה-LLMs כדי לפתוח את יכולות ההבנה הטקסטואליות שלהם.
תקציר מקורי באנגליתarXiv:2404.06349v3 Announce Type: replace Abstract: The ability to understand causality significantly impacts the competence of large language models (LLMs) in output explanation and counterfactual reasoning, as causality reveals the underlying data distribution. However, the lack of a comprehensive benchmark currently limits the evaluation of LLMs' causal learning capabilities. To fill this gap, this paper develops CausalBN-Bench based on data from the causal research community, enabling comparative evaluations of LLMs against traditional causal learning algorithms. To provide a comprehensive investigation, we offer three tasks of varying difficulties, including correlation, causal skeleton, and causality identification. Evaluations of 19 leading LLMs reveal that, while closed-source LLMs
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית