יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ORCA: ציד כשלי תכונה בדיפוזיה לתמונה

ORCA: Hunting Compositional Failures in Text-to-Image Diffusion
דיפוזיות לתמונה נכשלות בפני תבניות תכונה. ORCA מאינה את הלטנט של דיפוזיון-טרנספורמר עם יעד נמוך-דרגה.
תקציר מקורי באנגליתarXiv:2610.09841v1 Announce Type: cross Abstract: Text-to-image diffusion models fail predictably on compositional prompts: attributes bind to the wrong objects, spatial relations invert, and multi-object scenes lose count. Recent architectures already augment CLIP with a T5 encoder precisely because CLIP's contrastive embedding loses compositional structure, yet these failures persist. We argue the binding problem is therefore not one of missing information but of misaligned information: a text encoder preserves compositional structure, but in a representation space shaped by language modelling rather than vision, and the denoising objective does not directly reward aligning the two. We show this correspondence can be supplied as an explicit training signal, that the relevant cross-modal
קרא במקור המקורי