כתבה
arXiv cs.CL ·
Benchmarking Candidate Coverage in Typed Decision Models
תקציר מקורי באנגליתarXiv:2610.03387v1 Announce Type: cross Abstract: Typed decision models return choices or distributions over answer options supplied at request time. Accuracy with complete options does not establish whether a model recognizes that a reference answer is missing or avoids rejecting valid candidates. We present a paired candidate-coverage benchmark protocol and an initial evaluation of Laya and Jev across AG News, DBpedia, Emotion, and TREC. The models receive identical frozen texts and requests: 300 calibration and 589 test texts yield 23,932 predictions per model. Present/absent pairs match ordinary candidate count, and name variants preserve descriptions, members, and order. Native rejection behavior differs sharply: at five TREC candidates with natural names, Laya detects 97.2% of missin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית