יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

חשיבה מחוץ לקופסה

Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively?
מחקר חדש מציג את Box$^2$-Bench, כלי לבדיקת יכולתם של מודלי שפה להיעזר בהדרכה חיצונית באופן סלקטיבי. המחקר מראה כי מודלים מתקדמים יכולים ללמוד להיעזר בהדרכה מהימנה תוך כדי התעלמות מהדרכה לא מהימנה.
תקציר מקורי באנגליתarXiv:2609.39578v2 Announce Type: replace Abstract: Agent harnesses often improve language models with human-designed workflows, but as models grow more capable, unreliable guidance can increasingly constrain their execution. We call the ability to benefit from useful guidance while overriding unreliable guidance thinking outside the box. We introduce Box$^2$-Bench, which holds the model and task fixed while varying workflow reliability to isolate how models regulate their reliance on guidance. On Box$^2$-Bench, frontier models often benefit from reliable guidance but remain vulnerable when it is misleading or becomes unreliable. To test whether this capability can be learned, we train two open-weight models using bad workflows, reserving good workflows for evaluation. We explore two compl
קרא במקור המקורי