כתבה
arXiv cs.AI ·
מודל ראייה-שפה מאוחד לייצור דיווחי PET/CT, שאלות ויזואליות וסגמנטציה של רפסור
A Unified Vision-Language Model for PSMA PET/CT Report Generation, Visual Question Answering, and Lesion Segmentation
מודל ראייה-שפה מאוחד לייצור דיווחי PET/CT, שאלות ויזואליות וסגמנטציה של רפסור. המודל משתמש בארכיטקטורה של LLaVA וכולל רשת תפיסה ופרויקציה של MLP-Mixer. המודל עובד על 5,747 דגימות PET/CT והציג תוצאות טובות יותר מול PET2REP ובסיס קפדוארי.
תקציר מקורי באנגליתarXiv:2609.15603v1 Announce Type: cross Abstract: Accurate PSMA PET/CT interpretation is central to prostate cancer management, yet existing PET/CT AI models typically address isolated tasks. We propose a unified PSMA PET/CT vision-language model for report generation, visual question answering, and lesion segmentation. The framework adopts an LLaVA-style architecture, comprising a PET/CT vision encoder, an MLP-Mixer projection module, a LoRA-tuned large language model, and a 3D segmentation branch. Training followed a four-stage strategy: vision encoder pretraining, projection-layer alignment, VLM fine-tuning, and final multitask tuning. Language tasks used 5,747 PSMA PET/CT datasets with paired reports, while segmentation used the PSMA subset of AutoPET. The model outperformed PET2REP an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית