כתבה
arXiv cs.AI ·
MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
תקציר מקורי באנגליתarXiv:2609.04947v2 Announce Type: replace-cross Abstract: Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT). However, excessively long reasoning trajectories incur substantial computational costs and significant KV-cache pressure. Existing CoT compression and alignment paradigms mainly rely on static rules or single-dimensional preferences, lacking fine-grained cross-modal constraints; as a result, they are prone to inducing visual laziness and hallucinatory reasoning. To address these issues, we propose Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples. In the compression sta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית