יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

כאשר תקיפה עוזרת: מעקב ושליטה בשרשרת-החשבון במדיניות תצוגה-שפה-פעולה

When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies
במאמר זה, נחקר כיצד ניתן לשמור על תקיפה ולשלוט בשרשרת-החשבון במדיניות תצוגה-שפה-פעולה. המחברים מציגים דגם חדש לעקיבה ושליטה, ומדגימים את יעילותו במספר תחומים.
תקציר מקורי באנגליתarXiv:2610.00601v1 Announce Type: cross Abstract: Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning corr
קרא במקור המקורי