כתבה
arXiv cs.LG ·
הפחתת תווים: הסבה עצמית על-מדיניות להקטנת תווים חד-מדיניים
Fewer Tokens, More Self-Teaching: On-Policy Self-Distillation for Extreme Visual Token Reduction
הפחתת תווים: הסבה עצמית על-מדיניות להקטנת תווים חד-מדיניים. LT-OPD, פרקטיקה חדשה להקטנת תווים, משפרת את הביצועים של MLLMs תחת תקציב תווים נמוך.
תקציר מקורי באנגליתarXiv:2609.32353v2 Announce Type: replace-cross Abstract: Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student roll
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית