כתבה
arXiv cs.AI ·
UrbanVLA: מודל ראייה-שפה-פעולה למיקרו-ניידות עירונית
UrbanVLA: A Vision-Language-Action Model for Urban Micromobility
UrbanVLA הוא מודל ראייה-שפה-פעולה לניידות עירונית. הוא משלב נתונים חזותיים והנחיות קוליות כדי לנווט רובוטים בסביבות עירוניות. המודל עובר אימון בשני שלבים: אימון מונחה ואימון חיזוקי.
תקציר מקורי באנגליתarXiv:2510.23576v2 Announce Type: replace-cross Abstract: Urban micromobility applications, such as delivery robots, demand reliable navigation across large-scale urban environments while following long-horizon route instructions. This task is particularly challenging due to the dynamic and unstructured nature of real-world city areas, yet most existing navigation methods remain tailored to short-scale and controllable scenarios. Effective urban micromobility requires two complementary levels of navigation skills: low-level capabilities such as point-goal reaching and obstacle avoidance, and high-level capabilities, such as route-visual alignment. To this end, we propose UrbanVLA, a route-conditioned Vision-Language-Action (VLA) framework designed for scalable urban navigation. Our method
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית