כתבה
arXiv cs.CL ·
Ranking-Aware Prompt Optimization for Multimodal Clinical Diagnosis
תקציר מקורי באנגליתarXiv:2609.40361v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) are rapidly advancing clinical diagnosis, yet their adaptation pipelines remain anchored to accuracy-based objectives. Clinical data are heavily class-imbalanced: a constant-majority predictor can score above 90% accuracy while being clinically useless. We therefore evaluate and optimize for AUROC, a threshold-free score that ranks positives above negatives and is invariant to class balance. We focus on prompt optimization in MLLMs. Reflective methods such as GEPA use a binary scores matrix with one row per evaluation instance and one column per candidate prompt; cells record per-instance correctness, so the column average is accuracy and drives candidate selection. We introduce pair-level Pareto pro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית