כתבה
arXiv cs.AI ·
IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
תקציר מקורי באנגליתarXiv:2609.13725v1 Announce Type: new Abstract: An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, chang
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית