כתבה
arXiv cs.LG ·
AdaVLA: האצת מודלים חזות-שפה-פעולה
AdaVLA: Adaptive Step Flow Matching for Training-free Acceleration of Vision-Language-Action Models
AdaVLA היא שיטה חדשה להאצת מודלים חזות-שפה-פעולה. היא מאפשרת האצה משמעותית של המודלים, בלי צורך באימון מחדש. השיטה נבדקה על מודלים שונים, כולל SmolVLA, והראתה תוצאים מבטיחים.
תקציר מקורי באנגליתarXiv:2608.29208v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) models, built upon Vision-Language Models (VLMs), have significantly enhanced robotic capabilities by leveraging internet-scale knowledge and multimodal reasoning. However, the intensive computational overhead of VLAs constrains on-device deployment, hindering real-time responses to environmental changes. While various acceleration techniques have been proposed, they often rely on fine-tuning or access to training datasets, which are frequently unavailable due to privacy and proprietary concerns. Moreover, although flow-matching-based VLAs have emerged as efficient alternatives to standard diffusion models, current acceleration efforts largely target VLM inference costs, failing to address the iterative
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית