כתבה
arXiv cs.CL ·
ראייה אינה עומס: גריסה חד-פעימית לפיענוח ספקולטיבי במודלים של ראייה-שפה
Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
GLANCE הוא אלגוריתם חד-פעימי לגריסת בלוקים במודלים של ראייה-שפה. הוא מאפשר פיענוח ספקולטיבי מהיר יותר ומדויק יותר. GLANCE קורא את ההקשר הרב-מודאלי פעם אחת, ומקצר את זמן הפיענוח.
תקציר מקורי באנגליתarXiv:2609.00355v2 Announce Type: replace-cross Abstract: Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית