יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CoVeR: גזירת טוקנים מבוססת כיסוי לתפיסה תלת-ממדית

CoVeR: Coverage-Based Token Pruning for Multi-View 3D Reasoning in VLMs
CoVeR הוא אלגוריתם חדש לגזירת טוקנים מיותרים במודלים רב-תצוגתיים. הוא משפר את ביצועי התפיסה התלת-ממדית ומקטין את כמות הטוקנים הנדרשים. CoVeR עובד עם מודלים שונים ומראה תוצאות טובות יותר מאלגוריתמים קודמים.
תקציר מקורי באנגליתarXiv:2609.08345v1 Announce Type: cross Abstract: Representing a 3D scene as multi-view images allows 2D VLMs to reason in 3D by reusing priors from pre-training, sidestepping the scarcity of annotated 3D data. However, it produces thousands of redundant visual tokens whose cost grows with every view. Existing visual token pruners fall into two families, each limited in the 3D multi-view setting. Learned importance methods rank tokens by attention or encoder features; because redundancy here is fundamentally spatial, they keep near-duplicate tokens from a few prominent regions and leave most of the scene unrepresented. Voxelization methods improve spatial coverage but cannot enforce an exact token budget and saturate as multi-view observations overlap in 3D, capping retention well below th
קרא במקור המקורי