יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

PACE: תזמון תקין-מדיניותי להחלטה עד-קצב

PACE: Policy-Native Adaptive Decision Timing for Long-Horizon Reasoning
מדלג חדש: פוליצי-נטיבי ולמד-מודלי להחלטה עד-קצב, המגדיר את עומק הביצוע באופן תלוי בהיסטוריה.
תקציר מקורי באנגליתarXiv:2605.09860v5 Announce Type: replace Abstract: Long-horizon reasoning requires deciding not only what actions to take, but how many to execute open-loop before replanning. This number, the execution depth, balances replanning cost against compounding execution errors. Most current systems either fix the execution depth as a hand-tuned scalar or adjust it at inference time with heuristic rules decoupled from the policy; we argue both can be suboptimal. In this work, we treat the execution depth as a learnable, history-conditioned variable of the policy itself, and propose PACE, a model-native VLM policy that jointly predicts what to execute and for how long under a hard decision budget. We evaluate PACE on four long-horizon environments spanning fully and partially observable settings
קרא במקור המקורי