יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

IBBench-Light: בדיקה זוגית של תגובות-מותאמות-למשימה להנחיות חיצוניות

IBBench-Light: A Paired Evaluation of Task-Conditioned Responses to External Directives
בדיקה זוגית של תגובות-מותאמות-למשימה להנחיות חיצוניות. IBBench-Light בודק שני שימושים נגד האותה רשומה. שנים עשר יסודים סמנטיים יוצרים 144 זוגות מתואמים לכל מודל. ארבעה דגמי הנחיות-מוצפנות יצרו 1,152 תגובות-גרסה-גרסאות. דיוק-זוגי-מדויק-מדויק (PECA) דורש ששני החברים יקיימו את חוזי-התוצאה שלהם.
תקציר מקורי באנגליתarXiv:2609.13725v2 Announce Type: replace Abstract: An external record may contain a procedure to apply or text to read, depending on the user's request. IBBench-Light tests both uses against the same record. Twelve semantic bases yield 144 matched pairs per model; four quantized instruction models produced 1,152 archived greedy responses. Paired exact-contract accuracy (PECA) requires both members to satisfy their output contracts. Qwen succeeds on 132 execute and 109 process prompts, but only 97 complete pairs, showing what marginal averages omit. We audit literal-target exposure and case normalization, then add 1,722 logged CPU generations to test directive-absent controls, twelve additional semantic bases, within-base wording changes, and generation stopping. In the pinned Phi rerun, c
קרא במקור המקורי