יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Slot2Text: תפריט 2-Text: תפריט תצלומי-מרכזי לתפריט תצלומי-מרכזי לתפריט תצלומי-מרכזי

Slot2Text: Object-Centric Visual Tokenization for Efficient and Spatially Traceable Surgical MLLMs
Slot2Text: פיתוח של MLLM רפואי יעיל ועם תצלומי-מרכזי של תפריט תצלומי-מרכזי. המאמר עוסק בהצגת Slot2Text, דגם MLLM רפואי דו-מודי שמחליף תפריט תצלומי-מרכזי צפוף של תצלומי-מרכזי עם קבוצה קצרה של אזורים מקודדים כלטנטים. Slot2Text-Fast משתמש בקידוד האזור הקדמי לטפל בשאלות רפואיות. Slot2Text-Reason גם מזהה ומקודד אזורים רלוונטיים להסבר, ומקשר את הפלט של השפה לסמני-טקסט של האזור, המסכות או האזורים. הניסויים על מספר תשתיות שאלות-תצלומי-מרכזי ותשתיות-תצלומי-מרכזי-קריאה מראים ש- Slot2Text-Fast הוא תחרותי עם הבסיס המקורי של המצביא בעלות נמוכה יותר, ומקטין את צריכת הטקסט המוחלטת ב-91.8% ואת הקידוד הקדמי של התצלומי-מרכזי מ-1,295 ל-47 (הפחתה של 96.4%). Slot2Text-Reason מסחר על טקסט נוסף ועלייה בעלות לזהויות-אזור חדיד, מיקודים ולראיה-מרחבית-שפתית.
תקציר מקורי באנגליתarXiv:2608.01473v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLM) for surgical scene understanding typically inject hundreds of dense visual tokens into a language model, leading to costly inference and limited spatial traceability for generated answers. We present Slot2Text, a dual-mode surgical MLLM that replaces dense representations of visual input with a compact set of regions encoded as slot latents. Instead of relying on contrastive alignment of the visual encoder with language, Slot2Text groups self-supervised vision features into a few regions--slots that are consumed by the language model as area-labeled visual tokens. Slot2Text-Fast uses the slot prefix to answer surgical questions. Slot2Text-Reason also identifies and locates areas relevant for r
קרא במקור המקורי