כתבה
arXiv cs.AI ·
ראייה אינה עומס: גריעת בלוקים בסבב אחד לפענוח ספקולטיבי איבוד-חסר במודלים של שפה-ראייה
Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
GLANCE הוא מנגנון גריעת בלוקים בסבב אחד המשפר את קצב הפענוח הספקולטיבי במודלים של שפה-ראייה. הוא מאפשר קריאה אחת של ההקשר הרב-מודאלי, ללא קומפרסיה או הסתרה של התמונה. GLANCE מגיע לקצב פענוח מהיר פי 3.05 מאשר פענוח אוטורגרסיבי.
תקציר מקורי באנגליתarXiv:2609.00355v3 Announce Type: replace Abstract: Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree in one
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית