כתבה
arXiv cs.AI ·
COMiT: Learning Structured Visual Tokens through Sequential Communication
תקציר מקורי באנגליתarXiv:2602.20731v2 Announce Type: replace-cross Abstract: Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the existing token sequence. After several iterations, the resulting message conditions a flow-matching decoder that reconstructs the complete image. The encoder and decoder are implemented within a single tra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית