כתבה
arXiv cs.AI ·
CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
תקציר מקורי באנגליתarXiv:2605.12938v2 Announce Type: replace-cross Abstract: Video world models should predict future appearance in a way that remains consistent with 3D scene structure, camera motion, and lens geometry. Existing attention-level camera encodings, however, either describe each token only by its viewing ray---without locating scene content along that ray---or assume pinhole projection, limiting camera control under wide-angle and fisheye lenses. We introduce Curved Ray Expectation Positional Encoding (CRePE), which represents each image token as a depth-aware distribution along its Unified Camera Model (UCM) ray and integrates the expected rotary positional phasor along the curved path this distribution traces when projected into each query view. CRePE is realized through a lightweight Geometr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית