יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מדליות ופעולות: בניית דגמי פעולה יעילים על תיאורי נבואי חזותי

V-JEPA Policy: Building Effective World-Action Models on Predictive Visual Latents
במאמר זה, נראה שדגמי פעולה יעילים יכולים להיבנות על תיאורי נבואי חזותי. המחברים מציגים פרקטיקה פשוטה שבונה דגם פעולה על המרחב הלטנטי של עורך V-JEPA 2.1. הם מציעים נוסחאות פשוטות שבונות נבואי חזותי ומומחה פעולה זרימה-מתאים, שנלמדים יחד מן הראש בשלב יורד, עם תפקידי קו-מפתח-ערך של הנבואי המודע לעתיד. הם מציגים תוצאות תחרותיות עם דגמי WAM ובסיסי תכנות-שפה-פעולה מייצגים, ומציעים תוצאות טובות יותר כאשר הם מכשירים את הנבואי ללמוד על זוגות וידאו-הוראות ללא תווי-פעולה.
תקציר מקורי באנגליתarXiv:2609.37250v1 Announce Type: cross Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrained at scale, a prominent line of recent WAMs inherits both predictive knowledge and the models in which it was learned. We ask whether a predictive visual latent space induced by large-scale predictive pretraining can instead provide a sufficient foundation for effective WAM learning without inheriting a complete pretrained visual generative model. To answer this question, we introduce V-JEPA Policy, a simple framework that builds a WAM on the latent space of a frozen V-JEPA 2.1 encoder. An instruction-conditioned future-latent predictor and a flow-matching action expert are jointly learned fr
קרא במקור המקורי