יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

GroundAnything: יישום מקביל עם עגינה חזותית מדויקת

GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
GroundAnything הוא מודל יסוד לעגינה חזותית שמאפשר פענוח מקביל מהיר עם מיקום מדויק. המודל משלב אימון מוקדם, המרה ישירה ואימון עם פרסים. הוא מתעלה על מודלים אחרים בביצועים.
תקציר מקורי באנגליתarXiv:2609.39600v1 Announce Type: cross Abstract: Autoregressive (AR) grounding models serialize spatial predictions, introducing sequential latency and imposing a causal order on output tokens. We view grounding as visual evidence extraction: objects, locations, and spatial relations are jointly constrained by the image and query, yet their dependencies do not imply an intrinsic left-to-right generation order. This distinction makes bidirectional diffusion a natural fit, allowing spatial hypotheses to emerge in parallel and be jointly refined through iterative denoising. We introduce GroundAnything, a 4B-parameter grounding foundation model that reconciles fast parallel decoding with precise localization through blockwise denoising. Training combines grounding pretraining from public data
קרא במקור המקורי