כתבה
arXiv cs.LG ·
ELSA3D: Elastic Semantic Anchoring for Unified 3D Understanding and Generation
תקציר מקורי באנגליתarXiv:2607.06565v2 Announce Type: replace-cross Abstract: Unified 3D foundation models aspire to generate 3D assets and reason about them in language within a single backbone, but their text-3D interaction remains largely implicit. Existing methods concatenate text and 3D tokens into a flat sequence and rely on self-attention, collapsing coarse structural cues and fine geometric details into one undifferentiated representation. We introduce ELSA3D, a unified 3D model that addresses this with elastic semantic anchoring, structuring language and geometric reasoning jointly along matched abstraction scales. ELSA3D represents geometry with a scale-aware octree tokenizer and introduces Anchor Tokens, sparse cross-modal units that select semantic cues, route them to the most relevant 3D scale, r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית