כתבה
arXiv cs.AI ·
UniEvo-VL: מתכון אימון עצמי לשיפור מודלים רב-מודאליים
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement
UniEvo-VL הוא מתכון אימון עצמי למודלים רב-מודאליים. הוא משתמש בהדרכה עצמית ומשפר את יכולות היצירה של המודלים. הניסויים הראו שיפור ביכולות היצירה של Qwen-image-2512.
תקציר מקורי באנגליתarXiv:2609.38721v1 Announce Type: new Abstract: Modern multimodal models bring generation and understanding into a single unified system, which enables them to provide and learn from their own feedback. Motivated by this unified capacity, we introduce UniEvo-VL, a self-evolving framework for multimodal models to learn from this constructive self-correction feedback during test-time compute. Instead of relying on a separate, often larger, teacher, we leverage their self-critiques as privileged information and ask a single multimodal model to act as both teacher and student with different contexts. The student only sees the vanilla question, while the teacher conditions on the privileged critique. Then training minimizes the per-state divergence between their denoising diffusion distribution
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית