יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ClinFusion: מודל LLM רב-מודאלי להבנה רפואית שלמה

ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding
ClinFusion הוא מודל LLM רב-מודאלי שמאפשר הבנה רפואית שלמה. המודל משלב יכולות חזותיות וטקסטואליות לצורך ניתוח תמונות רפואיות ויצירת דוחות רפואיים. ClinFusion עולה על מודלים אחרים כגון GPT-5.2 ו-Gemini-3-Flash בביצועים.
תקציר מקורי באנגליתarXiv:2607.24743v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) hold immense potential to revolutionize clinical practice, yet deploying them in the medical domain is fundamentally a vision-centric challenge: models must absorb knowledge from heterogeneous 2D and 3D medical images, and evaluation protocols must align with radiologists' clinical practice and provide an accurate, fine-grained and factualness-driven assessment. In this paper, we introduce ClinFusion, a vision-centric MLLM designed for holistic medical understanding that systematically addresses these limitations. We propose a compositional and cascaded vision encoder architecture featuring a Cascade Spatial-Aware Locality Fusion operator that unifies diverse 2D and native 3D medical image understand
קרא במקור המקורי