כתבה
arXiv cs.CL ·
לדעת אינו לבחור: מה נוסף יש לאימות ספציפי מעבר לעדיפויות יוצרות
Knowing Is Not Choosing: What Explicit Verification Adds Beyond Generative Preference
אימות ספציפי משפר את דירוג השאלות ודיוק הרוב. נמצא כי דגמי Gemma, Qwen3 ו-Llama יוצרים תוצאות טובות יותר כאשר משתמשים באימות ספציפי.
תקציר מקורי באנגליתarXiv:2609.33142v2 Announce Type: replace Abstract: Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct candidate, ranking the available candidates, and selecting the final answer. Pre-generation readouts predict factual recall and which questions sampling will cover across three model families, but say little about whether an available correct answer will ultimately be selected. Explicit verification with $P(\mathrm{True})$ improves within-question ranking over mean log-likelihood in Gemma, Qwen3, and Llama, with AUROC gains of $0.08$--$0.12$. In a prospectively defined Gemma cohort, verification raises plurality accuracy by about $5$ points, and still gains about $2$ points over chat-template li
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית