יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

דגמים מהירים, ראיות ארוכות: בירור ובדיקה עצמית של דגמי החלטה מס' 1 להרשמות LLM

Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent Harnesses
בירור ובדיקה עצמית של דגמי החלטה מס' 1 להרשמות LLM. המחקר משווה שני דגמים: Laya ו-Jev, ומצא ש-Jev היה יותר נכון ב-9 מ-11 נקודות החלטה.
תקציר מקורי באנגליתarXiv:2610.02267v1 Announce Type: cross Abstract: Agent harnesses make many small, typed decisions per task: which model to call, which tool to use, whether retrieved text is relevant, whether an input carries an injection. System-1 decision models answer such questions in a single forward pass with class probabilities, promising large cost and latency savings over LLM calls. We present a paired evaluation of an open-weight (Laya) and a hosted (Jev) System-1 model on 11 agent decision points built from 18 public sources: 7,283 base cases plus 6,640 robustness variants, with byte-identical inputs, paired tests, and cross-hardware and cross-day reproducibility checks. Jev is significantly more accurate on 9 of 11 decision points (+10.8 to +46.0 pp). Neither model beats chance on zero-shot mo
קרא במקור המקורי