יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

רובוט מצטער, אדם מאושר: מודלים ויז'ואלי-לשוניים קוראים רק אחד משני שכבות טיפוגרפיות קריאות

Sorry Robot, Happy Human: Vision-Language Models Read Only One of Two Legible Typographic Layers
מודלים ויז'ואלי-לשוניים מתקשים לקרוא טקסט עם שכבות טיפוגרפיות מרובות. נוצר מאגר נתונים DecoyBench עם 300 תמונות, ונבדקו 6 מודלים סגורים. התוצאות הראו שבני אדם יכולים לקרוא את שתי השכבות, אך המודלים קוראים רק את השכבה הראשונה.
תקציר מקורי באנגליתarXiv:2609.31403v1 Announce Type: new Abstract: Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBench dataset was created using the Decoy Font method. The dataset consists of 300 images, each containing text with sharp contour lines superimposed on another text with soft shading. Six recent closed-source models from three different model families were evaluated using this dataset under two different prompting conditions (naive and guided) and at two different resolutions ($512\times512$ and $64\times64$). A validation study showed that human participants could read both text layers with high accuracy. In contrast
קרא במקור המקורי