יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מה גורם לכלליות תחבירית בדגמי ייצור ראייתיים?

What Drives Compositional Generalization in Visual Generative Models? The Importance of Continuous Training Objectives
חוקרים חקרו את כלליות תחבירית בדגמי ייצור ראייתיים. הם גילו שאפקטיביות תלויה בסוג האובייקטיב המשמש.
תקציר מקורי באנגליתarXiv:2510.03075v4 Announce Type: replace-cross Abstract: Compositional generalization, the ability to generate novel combinations of known concepts, is a key ingredient for visual generative models. Yet, not all mechanisms that enable or inhibit it are fully understood. In this work, we conduct a systematic study of which design choices critically determine compositional generalization in image and video generation. By isolating independent design axes, we identify two key factors strongly associated with compositional success: (i) whether the training objective operates on a discrete or continuous distribution, and (ii) the completeness of conditioning information about constituent factors during training. We also show that relaxing the discrete loss with an auxiliary continuous latent o
קרא במקור המקורי