יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ראייה אינה עומס: גריעת בלוקים בסבב אחד לפיענוח ספקולטיבי איבוד-פונקציות במודלים של ראייה-שפה

Vision Is Not Overhead: One-Pass Block Drafting for Lossless Speculative Decoding in Vision-Language Models
GLANCE הוא אלגוריתם חדש לגריעת בלוקים בסבב אחד, המאפשר פיענוח ספקולטיבי מהיר יותר במודלים של ראייה-שפה. הוא משתמש בראש בלוק-דיפוזיה כדי ליצור בלוקים שלמים בסבב אחד, ומאפשר למודל לקרוא את ההקשר הרב-מודאלי פעם אחת. GLANCE מהיר פי 3.05 מאשר פיענוח אוטורגרסיבי, ועוקף את המודל EAGLE3-VL במשימות מובנות.
תקציר מקורי באנגליתarXiv:2609.00355v3 Announce Type: replace-cross Abstract: Speculative decoding accelerates generation without changing its output, but on vision-language models (VLMs) a self-reinforcing cycle holds it back. Because an autoregressive drafter pays a sequential pass for each drafted token, it must stay small and can ill afford to attend to the image at each pass. Prior work therefore compresses or hides the image, leaving the drafter weakest on the text the image determines. We present GLANCE, a one-pass block drafter that breaks this cycle on an unmodified VLM target. Its block-diffusion head drafts a whole block in one forward pass over the target's already fused vision-language states, reading the multimodal context once, however deep the draft. The target verifies a wide candidate tree i
קרא במקור המקורי