יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Dynin-Robotics: מודל אחד לראייה, שפה ופעולה

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model
Dynin-Robotics הוא מודל אומנימודלי שמאפשר חיזוי מטרות ויזואליות ושינויים בסצנה. המודל משלב יכולות שונות, כולל חיזוי פעולות, חיזוי מצבים עתידיים ושיחזור הוראות. המודל הוכשר על מיליוני טראג'קטוריות והותאם לתחומים שונים.
תקציר מקורי באנגליתarXiv:2609.13053v1 Announce Type: cross Abstract: Visual goal and dynamics prediction can provide language-conditioned robot policies with both a target outcome and a representation of action-dependent scene changes. We bring these predictions into action generation and selection through a shared trajectory model. Dynin-Robotics implements this formulation on Dynin-Omni, an omnimodal masked-diffusion backbone, representing language, visual observations, goals, and actions as discrete tokens. By varying conditioning and target spans, the same model learns action prediction, action-conditioned next-observation prediction, terminal goal-state prediction, and trajectory-to-instruction reconstruction. These interfaces support test-time scaling through goal prediction, action-candidate evaluatio
קרא במקור המקורי