כתבה
arXiv cs.CL ·
SpikingVLA: דגמי ראייה-שפה-פעולה ספיקים
SpikingVLA: Asynchronous Spiking Vision-Language-Action Models
SpikingVLA: דגמי ראייה-שפה-פעולה ספיקים שמשפרים את הערכיות והקצב
תקציר מקורי באנגליתarXiv:2610.09710v1 Announce Type: new Abstract: ANN-to-SNN conversion offers a practical route toward energy-efficient spiking Vision-Language-Action (VLA) models by bypassing the substantial cost of training large-scale SNNs from scratch. However, existing methods often require many timesteps to maintain competitive performance, resulting in substantial inference latency for real-time VLA deployment. To address this challenge, we introduce SpikingVLA, an ANN-to-SNN conversion framework that enables accurate and low-latency spiking VLA inference. Specifically, we propose a Dendritic Integrate-and-Fire (DIF) neuron that alleviates channel-wise activation outliers through dendritic mixing and adaptive somatic firing, enabling accurate ANN-to-SNN conversion with fewer timesteps. Building on D
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית