יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

COMiT: Learning Structured Visual Tokens through Sequential Communication

תקציר מקורי באנגליתarXiv:2602.20731v2 Announce Type: replace-cross Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the existing token sequence. After several iterations, the resulting message conditions a flow-matching decoder that reconstructs the complete image. The encoder and decoder are implemented within a single tra
קרא במקור המקורי