כתבה
arXiv cs.AI ·
Time-Frequency Geometric Cross-Attention for Chunked Vision-Language-Action Models
תקציר מקורי באנגליתarXiv:2609.09925v1 Announce Type: new Abstract: Modern vision-language-action (VLA) policies predict a whole chunk of actions: one to two seconds of coordinated motion emitted in a single forward pass. Yet an action chunk is essentially a short multivariate trajectory, but inside these models it is a sequence of generic per-timestep hidden tokens decoded by a linear head. This under-serves two motion structures. First, frequency: a chunk superimposes a smooth global trend and fine corrective motion across time scales, and a single token entangles them. Second, cross-phase geometry: motions of different phases (reach, contact, grasp adjustment, settling) unfold along very different, near-orthogonal directions in representation space, yet are tightly related for the task and arise across the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית