כתבה
arXiv cs.CL ·
CruxBench: תקן לגילוי מידע
CruxBench: A Benchmark of Information Discovery
CruxBench הוא תקן לבדיקת יכולת הגילוי של דגמי שפה גדולים. הוא מדד את יכולת הדגמים לגלות תשובות חשובות לשאלות קשות. CruxBench נחשב לתקן חדשני ומקיף, שמאפשר לבדוק יכולות של דגמי שפה בצורה יעילה ואמינה.
תקציר מקורי באנגליתarXiv:2609.35879v1 Announce Type: new Abstract: Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by construction
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית