יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

עבר לממדים: סכסוך גרדיאנטים ניתן-מיקום ומודולציה-מודעת-מיקום למודלים-מולטימודליים מאוחדים

Beyond Layers: Position-Resolved Gradient Conflict and Position-Aware Modulation for Unified Multimodal Models
מאמר זה חושף כי סכסוך בין הבנת תמונות ליצירת תמונות במודלים-מולטימודליים מאוחדים תלוי במיקום. המחברים הציגו מפה של סכסוך-גרדיאנטים-ניתן-מיקום והציעו מודולציה-מודעת-מיקום (PAM) להפחית את הסכסוך.
תקציר מקורי באנגליתarXiv:2609.38485v1 Announce Type: cross Abstract: Unified multimodal models (UMMs) train image understanding and autoregressive image generation on shared parameters, and the two objectives are known to interfere. Existing diagnoses and remedies operate at the resolution of layers or experts, measuring conflict per layer and resolving it by separating parameters. We argue that this resolution hides an orthogonal axis. Generation in a UMM is next-token prediction over a raster sequence of visual tokens whose roles vary systematically with position, so how strongly a generation gradient interferes with understanding should depend on where in the sequence it originates. We introduce a position-resolved interference map that attributes understanding-generation gradient conflict to visual-token
קרא במקור המקורי