יום חמישי, 30 ביולי 2026 LIVE
AI־INFO

כתבה MarkTechPost ·

Robbyant משיק LingBot-VA 2.0: מודל וידאו-פעולה קוזאלי

Ant Group’s Robbyant Unveils LingBot-VA 2.0: A Causal Video-Action Model Built Natively for Physical AI
Robbyant השיקה את LingBot-VA 2.0, מודל וידאו-פעולה קוזאלי לשליטה ברובוטים. המודל מאפשר שליטה טובה יותר בפעולות הרובוט ומבוסס על ארכיטקטורה חדשה.
תקציר מקורי באנגליתRobbyant, the embodied AI unit inside Ant Group, has released the LingBot-VA 2.0 .The first embodied-native foundation model. It describes a video-action foundation model for generalist robot manipulation. The research team pretrains the whole stack for embodiment instead of fine-tuning a video generator. What is LingBot-VA 2.0? Most video-action models reuse two components built for digital content creation. One is a reconstruction-oriented VAE. The other is a bidirectional video-diffusion backbone, with an action module attached. This creates three limitations. Pixel-reconstruction latents preserve appearance but carry limited physical structure. Iterative denoising over video tokens is too slow for closed-loop control. Generic video objectives never teach how actions reshape the world.
קרא במקור המקורי