כתבה
arXiv cs.AI ·
V-Engram: Trigger-Indexed External Memory for Modular Text-to-Image Personalization
תקציר מקורי באנגליתarXiv:2609.37198v1 Announce Type: new Abstract: Pretrained text-to-image models contain broad visual knowledge, yet they cannot reliably acquire or refine a specific visual identity from only a few references while preserving compositional control. Token-embedding methods are compact but often underfit identity, whereas adapter-based methods improve fidelity through persistent weight updates that can be costly to store and interfere when concepts are composed. We introduce V-Engram, a trigger-indexed external memory mechanism for Stable Diffusion 3.5. Each concept is assigned an explicit trigger that retrieves concept-specific memory, whose gated directions enter frozen text-encoder and MMDiT context states as relative residuals. Separating this memory from backbone adaptation enables prom
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית