כתבה
arXiv cs.AI ·
Less Decoder is More Encoder: Geometric Representation Learning from Novel View Synthesis
המאמר חוקר את תפקידה של Novel View Synthesis (NVS) בלמידת התנהגות גאומטרית. SNAP, תרגומן העצמי של NVS, מציג תצוגה גאומטרית טובה יותר.
תקציר מקורי באנגליתarXiv:2610.03717v1 Announce Type: cross Abstract: This paper examines the role of Novel View Synthesis (NVS) in geometric representation learning. In principle, NVS should reason about 3D scene structure, thereby enabling transferable multi-view geometric representations. Yet, existing encoder-based NVS methods yield poor representations. This is not because of a lack of supervisory signal, but rather due to inconspicuous architectural choices: \textit{spatially expressive decoders} that dilute representational capabilities of the scene encoder, and \textit{low-level pixel-space targets} that hinder feature learning. We present SNAP, a self-supervised encoder-decoder transformer that addresses both through a pose-conditioned local decoder and a latent-space reconstruction objective. SNAP i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית