יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

JEPA-Bisim: לימוד ייצוגים ויזואליים עמידים

JEPA-Bisim: Learning Robust Visual Representations for Planning with Joint-Embedding Predictive World Models
JEPA-Bisim משפר את העמידות בתכנון ויזואלי. המודל משתמש בארכיטקטורת JEPAs ומוסיף מקודד ביסימולציה. הוא מראה שיפור בעמידות לשינויים ויזואליים.
תקציר מקורי באנגליתarXiv:2602.18639v2 Announce Type: replace Abstract: World models learned from high-dimensional visual observations allow agents to make decisions and plan directly in latent space, avoiding pixel-level reconstruction. However, recent latent predictive architectures (JEPAs), including the DINO world model (DINO-WM), display a degradation in test time robustness due to their sensitivity to ``slow features". These include visual variations such as background changes and distractors that are irrelevant to the task being solved. We address this limitation by augmenting the predictive objective with a bisimulation encoder that enforces control-relevant state equivalence, mapping states with similar transition dynamics to nearby latent states while limiting contributions from slow features. We ev
קרא במקור המקורי