כתבה
arXiv cs.AI ·
Labels Override Definitions in Jev-Style Typed Decision Models
תקציר מקורי באנגליתarXiv:2610.02586v1 Announce Type: new Abstract: A typed decision model answers a fixed question about an input by returning a probability for each of several caller-defined options. Each option carries a short label and a written definition, which is where a developer states the rule the model should apply. Jev introduced this interface for routing, moderation and triage, open implementations followed, and the same operation occurs whenever a language model is used as a classifier by scoring label strings. We study the open implementations, whose weights we can inspect and patch, and ask whether the probability follows the definitions or the labels. A preference for the label we call option-label bias. Across four open-weight typed decision models, three ways of reading an answer from a Qw
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית