יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אל תפלג את הזנב: Action Upcycling להגברת קצב המדיניות

Don't Throw Away the Tail: Action Upcycling for Policy Acceleration
Action Upcycling מאריך את האזור התפעולי של המדיניות, ללא גישה למודל או דגימות נוספות. ניסויים נרחבים במשימות תפעוליות מומחים ואמיתיות הראו כי Action Upcycling מקטין את מספר קריאות המדיניות ב-1.2-1.7x, עם תפוקה זהה.
תקציר מקורי באנגליתarXiv:2609.34911v2 Announce Type: replace-cross Abstract: Modern robot policies predict a chunk of future actions from a single observation, execute only a prefix, and discard the rest before replanning. Choosing the length of this prefix, the execution horizon, poses a trade-off between reactivity and efficiency. A short horizon keeps the policy reactive to the environment, but requires frequent policy calls. Recent test-time methods adaptively select the horizon for each chunk, but they either read model internals, where the signal must be chosen for each architecture, or draw extra samples, which adds cost. We propose Action Upcycling, a training-free algorithm that reuses actions the policy would otherwise discard, without accessing model internals or drawing extra samples. We find tha
קרא במקור המקורי