כתבה
arXiv cs.AI ·
MCPO: אופטימיזציה של נקודות עדיפות חופפי-מודלי להקטנת רצפי חשיבה במודלי CoT
MCPO: Modality-Contrastive Preference Optimization for Multimodal Chain-of-Thought Compression
MCPO היא שיטה להקטנת רצפי חשיבה במודלי CoT, המשתמשת באופטימיזציה של נקודות עדיפות חופפי-מודלי. השיטה כוללת שני שלבים: הקטנת רצף והתאמה. בשלב ההקטנה, השיטה משתמשת באלגוריתם של גיבוש NCMI להקטנת רצף. בשלב ההתאמה, השיטה משתמשת באופטימיזציה של נקודות עדיפות חופפי-מודלי להתאמה של רצף. השיטה הוכיחה את עצמה במבחן על מודל Qwen3-VL-Thinking, והציגה תוצאות טובות בהקטנת רצף ובשיפור עדינות.
תקציר מקורי באנגליתarXiv:2609.04947v1 Announce Type: cross Abstract: Recently, multimodal large-scale reasoning models have demonstrated remarkable capabilities in solving complex tasks through long Chains-of-Thought (M-CoT). However, excessively long reasoning trajectories incur substantial computational costs and significant KV-cache pressure. Existing CoT compression and alignment paradigms mainly rely on static rules or single-dimensional preferences, lacking fine-grained cross-modal constraints; as a result, they are prone to inducing visual laziness and hallucinatory reasoning. To address these issues, we propose Modality-Contrastive Preference Optimization (MCPO), a highly sample-efficient two-stage length-compression method that requires fewer than 900 training samples. In the compression stage, we i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית