כתבה
arXiv cs.LG ·
NAC: Neural Action Codec for Vision-Language-Action Models
תקציר מקורי באנגליתarXiv:2606.21372v2 Announce Type: replace-cross Abstract: Vision-language-action (VLA) models rely on discrete action tokenizers to bridge continuous robot control and autoregressive sequence modeling, yet existing tokenizers often trade off between compression, latency, and downstream performance. We revisit this design through the lens of neural audio codecs - convolutional encoder-decoder architectures with residual vector quantization that serve as the standard front end for audio foundation models. Motivated by their success, we introduce the Neural Action Codec (NAC), which treats short robot action trajectories as multi-channel 1D signals and compresses them using a multi-scale RVQGAN architecture. With adaptations to the action representation, compression rate, and reconstruction o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית