כתבה
arXiv cs.AI ·
Benchmarking candidate coverage and rejection policy transfer in typed decision models
תקציר מקורי באנגליתarXiv:2610.03387v2 Announce Type: replace Abstract: Rejection policies must remain useful as candidate sets and tasks change. We compare Laya, Jev and Qwen2.5-7B-Instruct using public reference labels, testing Laya/Jev policy transfer at equal calibration budgets and all three models on artificial omission, natural retrieval misses and public out-of-scope queries. Source calibration often fails to preserve the target operating point. A Jev policy calibrated on DBpedia rejects 69.3% of covered Emotion test inputs, while an Emotion policy loses detection entirely. Retrieval exposes a different tradeoff: with ten intent candidates, Laya detects 99.0% of out-of-scope queries but rejects 48.8% of covered queries. Separating missing-answer sources reveals these costs alongside retrieval coverage
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית