כתבה
arXiv cs.CL ·
Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
תקציר מקורי באנגליתarXiv:2605.30984v2 Announce Type: replace-cross Abstract: Modern 3D medical vision-language models (VLMs) can generate fluent radiology-style text while exhibit critically low pathology detection and output diversity, collapsing to generic templates that under-report rare yet critical findings. We identify this failure mode as Template Collapse. This failure stems from the unique constraints of 3D medical imaging, e.g., limited data, severe label imbalance, and weak signals from volumetric encoders. Under these constraints, text-generation objectives encourage shortcut learning and fluent but weakly grounded reports. We systematically diagnose the Template Collapse through clinical fidelity, output diversity, normal-template bias, and rare-finding survival. To mitigate it, we propose CLarG
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית