כתבה
arXiv cs.CL ·
שאלות ארקטיות, תשובות חסרות
Arctic Questions, Missing Answers: A Dataset and Benchmark for LLM Abstention in Arctic Science
חוקרים פיתחו מאגר נתונים בשם ArcticQA, המאפשר לבדוק את היכולת של מודלים כמו Gemini, Claude ו-GPT להימנע מתשובות שגויות. המחקר בדק 8 מודלים ומצא תוצאות מעניינות.
תקציר מקורי באנגליתarXiv:2610.09446v1 Announce Type: new Abstract: Large language models (LLMs) should abstain from scientific multiple-choice questions when no option is valid, but frequent abstention alone does not demonstrate sensitivity to answer availability. We introduce ArcticQA, a dataset of 194 questions derived from primary Arctic research, with automated checks of answer support and distractor contradiction against source evidence. We further develop ArcticAbstain, a paired benchmark comparing answer-present and answer-absent conditions, with the correct answer replaced by a distractor in the latter and an explicit abstention option in both. We evaluate eight models from the Gemini, Claude, and ChatGPT families at high reasoning effort, with three trials per condition, yielding 9,312 recorded resp
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית