כתבה
arXiv cs.AI ·
MedVL-SAM2: דגם ראיית-שפה 3D מאוחד לתגובה מולטימודלית ולחילופין פרומפט-ניתוח
MedVL-SAM2: A unified 3D medical vision-language model for multimodal reasoning and prompt-driven segmentation
דגם ראיית-שפה 3D מאוחד לתגובה מולטימודלית ולחילופין פרומפט-ניתוח. הדגם, MedVL-SAM2, משלב ראייה-שפה 3D עם ניתוח ופרומפט-ניתוח, ומציג תוצאות טובות בתחומי דיווח, תשובות לשאלות ויזואליות וניתוח 3D.
תקציר מקורי באנגליתarXiv:2601.09879v2 Announce Type: replace-cross Abstract: Recent progress in medical vision-language models (VLMs) has achieved strong performance on image-level text-centric tasks such as report generation and visual question answering (VQA). However, achieving fine-grained visual grounding and volumetric spatial reasoning in 3D medical VLMs remains challenging, particularly when aiming to unify these capabilities within a single, generalizable framework. To address this challenge, we proposed MedVL-SAM2, a unified 3D medical multimodal model that concurrently supports report generation, VQA, and multi-paradigm segmentation, including semantic, referring, and interactive segmentation. MedVL-SAM2 integrates image-level reasoning and pixel-level perception through a cohesive architecture ta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית