כתבה
arXiv cs.LG ·
FiberTune: Preserving Action-Fiber Visual Residuals in Vision-Language-Action Fine-Tuning
תקציר מקורי באנגליתarXiv:2606.08653v3 Announce Type: replace-cross Abstract: Action-supervised fine-tuning of vision-language-action (VLA) policies fits demonstrations effectively but constrains only the directions that change predicted actions, leaving visual structure consistent across action-equivalent states free to collapse. We formalize this as residual visual collapse along local action fibers and propose FiberTune, a training-time objective that preserves teacher-structured visual residuals without adding inference-time overhead. FiberTune uses an online action probe to estimate action-predictive feature directions, filters them from intermediate visual-token representations, and aligns the resulting probe-filtered residuals to a frozen visual teacher while regularizing their effective rank. Under id
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית