יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MoLE: מערכת של אקספרטים רצפיים לתפיסה חזותית משלימה

MoLE: Mixture of Latent Experts for Complementary Visual Reasoning
MoLE היא מערכת של אקספרטים רצפיים שמטרתה לשפר את תפיסת החזות המשלימה. המערכת מספקת תפיסה רצפית של ראיות חזותיות, ומאפשרת לאקספרטים לגלות ראיות חזותיות שונות.
תקציר מקורי באנגליתarXiv:2610.01917v1 Announce Type: cross Abstract: Latent visual reasoning equips vision--language models with continuous intermediate states that can process visual evidence without explicit textual reasoning traces or repeated image operations. However, existing methods often allow multiple latent tokens to access the same visual evidence through shared value projections, providing no mechanism for them to extract complementary visual information; simply increasing the latent budget can therefore yield redundant latent representations. We argue that effective latent reasoning should encourage different latent tokens to extract complementary visual information, and thereby act as specialized visual experts. Based on this insight, we propose MoLE, a Mixture of Latent Experts framework that
קרא במקור המקורי