כתבה
arXiv cs.AI ·
GeoLatent: מודל לתפיסת מרחב 3D
GeoLatent: Geometry-Guided Latent Structuring with Routed Optimization for 3D Reasoning
מודל GeoLatent משלב גיאומטריה ואופטימיזציה לתפיסת מרחב 3D מתמונות 2D. המודל משפר את הדיוק בתפיסת מרחב ומתאים למשימות כגון SPAR-Bench ו-SPBench.
תקציר מקורי באנגליתarXiv:2610.02091v1 Announce Type: cross Abstract: Despite progress in vision-language models, 3D spatial reasoning from 2D images remains challenging. Text-based methods describe intermediate geometry with discrete tokens, limiting fidelity for continuous spatial relations. Continuous latents offer richer representations, but a single latent type does not explicitly separate the cues needed across spatial tasks. Decomposed spatial latents address this by representing position, direction, and global geometry separately under geometric supervision. Yet the geometry representation can still collapse toward one dominant direction, and unrestricted attention can leave the latents underused during answer learning. We introduce GeoLatent, combining Common--Residual Geometry Alignment (CR-GEO) wit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית