כתבה
arXiv cs.AI ·
CruxBench: בנצ'מרק לגילוי מידע
CruxBench: A Benchmark of Information Discovery
CruxBench הוא בנצ'מרק להערכת יכולתם של מודלים גדולים לשפה (LLM) לגלות מידע. הוא בודק את היכולת לפרק שאלות מורכבות לשאלות משנה, ולהעריך את ערך המידע שלהן. הבנצ'מרק נותן ציון לפי ערך המידע (VOI) של השאלות המוצעות על ידי המודל.
תקציר מקורי באנגליתarXiv:2609.35879v1 Announce Type: cross Abstract: Benchmarks for large language models (LLMs) typically evaluate the accuracy of answers against fixed reference labels. But a central step in many complex real-world tasks is identifying which questions are worth asking in the first place: decomposing a difficult problem into subquestions -- which we call cruxes -- whose answers provide key steps on the path toward solving the target problem. To evaluate this capability of information discovery, we introduce CruxBench, a benchmark that grades LLM-generated questions by their Value of Information (VOI): how much a model-proposed crux updates beliefs about a target forecasting question. CruxBench enjoys a rare combination of three key properties: it is (1) contamination-resistant by constructi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית