יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

תרגום ויזואלי-לשוני משולב: חקירה ניהולית

Cross-Task Generalization Between Understanding and Generation in Unified Vision-Language Models: A Controlled Study
מודלי תרגום ויזואלי-לשוני משולבים משפרים שני תחומי ידע: הבנה ויצירה.
תקציר מקורי באנגליתarXiv:2505.23043v2 Announce Type: replace-cross Abstract: Unified vision-language models (VLMs) aim to support both visual understanding and generation within a single framework, but it remains unclear when mixed training benefits both capabilities and when it introduces conflicts. This paper presents a controlled empirical study of cross-task generalization between understanding and generation in unified VLMs. We construct two controllable image-text benchmarks, SmartWatch and modified CelebA, with paired VQA, captioning, and text-to-image generation tasks, and evaluate multiple LLM-based unified architectures built from SigLIP and VQ-VAE visual spaces. Our experiments show that mixed understanding-generation training can improve both tasks over task-specific training, but the benefit dep
קרא במקור המקורי