כתבה
arXiv cs.AI ·
CoVisco: Codec-Native Vision Encoder with Native Token Compression for Unified Image-Video Understanding
תקציר מקורי באנגליתarXiv:2609.39924v1 Announce Type: cross Abstract: Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making long-video understanding expensive for the vision encoder and the language model. Existing methods often compress visual tokens after dense encoding, creating a mismatch between the representation used during training and the compact interface required at deployment. We present CoVisco, a codec-native vision encoder with native token compression for unified image-video understanding. By combining codec-native input support with segmented attention, CoVisco can encode long visual inputs in a single forward pass without forming dense patch-to-patch interactions across all frames. Each tempo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית