כתבה
arXiv cs.AI ·
LensVLM: פיתוח תצוגה נבחרת של תצלומים נרחבים לייצוג גרפי של טקסט
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
LensVLM מאפשר למודלי VLM לסרוק תצלומים נרחבים ולהרחיב תצלומים רלוונטיים. התצלומים הנרחבים מאפשרים למודלים לקרוא טקסט ולבצע תפקודים שונים. המאמר מציג תיאור נרחב של פיתוח LensVLM והתצלומים הנרחבים שהוא מאפשר. התצלומים הנרחבים יכולים לשפר את תפקוד המודלים ולאפשר להם לבצע תפקודים שונים.
תקציר מקורי באנגליתarXiv:2605.07019v3 Announce Type: replace-cross Abstract: Vision Language Models (VLMs) offer the exciting possibility of processing text as rendered images, bypassing the need for tokenizing the text into long token sequences. Since VLM image encoders map fixed-size images to a fixed number of visual tokens, varying rendering resolution provides a fine-grained compression knob. However, accuracy deteriorates quickly as compression increases: characters shrink below the vision encoder's effective resolution, making them indistinguishable. To address this, we propose LensVLM, an inference framework and post-training recipe that enables VLMs to scan compressed images, then selectively expand only the relevant images to their uncompressed form via learned tools. Building on Qwen3.5-9B-Base, L
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית