יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

למידת תכונות נוף-בינאורתיות בעולם-מודלים דרך פערי פעולה-לטנט

Learning Visual Feature-Based World Models via Residual Latent Action
במאמר זה, המחברים מציגים חידוש בלמידת עולם-מודלים, המתמקד בלמידת תכונות נוף-בינאורתיות. הם מציגים תיאור חדש של פעולה-לטנט, המכונה Residual Latent Action (RLA), ומציגים עולם-מודל (RLA-WM) המשתמש ב-RLA. ה-RILA-WM מציג תוצאות טובות יותר מאשר עולם-מודלים קיימים, ומספק תיאור טוב יותר של תכונות נוף-בינאורתיות.
תקציר מקורי באנגליתarXiv:2605.07079v2 Announce Type: replace-cross Abstract: World models predict future transitions from observations and actions. Existing works predominantly focus on image generation only. Visual feature-based world models, on the other hand, predict future visual features instead of raw video pixels, offering a promising alternative that is more efficient and less prone to hallucination. However, current feature-based approaches rely on direct regression, which leads to blurry or collapsed predictions in complex interactions, while generative modeling in high-dimensional feature spaces still remains challenging. In this work, we discover that a new type of latent action representation, which we refer to as Residual Latent Action (RLA), can be easily learned from DINO residuals. We also s
קרא במקור המקורי