כתבה
arXiv cs.AI ·
DVD: Dynamic Vector Decoding for Efficient MLLM-based Perception
תקציר מקורי באנגליתarXiv:2610.12266v1 Announce Type: cross Abstract: Multimodal large language models have made remarkable progress in bridging vision and language, facilitating various perception tasks essential for human-machine interaction, robotics, and autonomous driving. However, existing MLLM-based perception methods predominantly rely on text-based coordinate representation, which suffers from excessive token overhead, or fixed-range quantization, which suffers from range and precision constraints, especially for 3D domains with unbounded spatial range and high localization accuracy requirements. To address these challenges, we propose a dynamic vector decoding method named DVD, which unifies the representation of 2D and 3D perception tasks. Specifically, we first transform diverse perceptual represe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית