כתבה
arXiv cs.AI ·
Cooperative Multi-Agent Vision-Language-Action Models via Reinforced Fine Tuning
תקציר מקורי באנגליתarXiv:2609.36588v1 Announce Type: cross Abstract: We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLAs are pretrained on large-scale single-agent data and therefore lack the fine-grained coordination skills required for inter-robot collaboration. Supervised fine-tuning (SFT) on multi-robot demonstrations partially bridges this gap, but its performance is bounded by the demonstration data and cannot improve from its own experience. We present a three-stage reinforced fine-tuning (RFT) pipeline for multi-agent VLAs. First, initialization-aware data collection sweeps over initial configurations and invokes human demonstrations only when the pretrained VLA repeatedly fails, yielding robustness to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית