יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

נקודות תורפה של MLLM

Where MLLMs Fail and Why: Causal Task Decomposition for Capability Failure Diagnosis
חוקרים הציגו שיטה לאבחון נקודות תורפה במודלים ללמידת מכונה. השיטה מסייעת להבין מדוע מודלים אלו נכשלים במשימות מסוימות. המחקר מראה כי אספקת נתונים נכונים יכולה לצמצם את שגיאות המודלים.
תקציר מקורי באנגליתarXiv:2609.38851v1 Announce Type: new Abstract: End-to-end accuracy on compositional tasks records how often MLLMs fail, but cannot distinguish whether a failure reflects an intrinsic deficit in the targeted capability or a cascading error from an upstream prerequisite. We propose a causal decomposition framework that isolates these two failure modes through controlled interventions on the prerequisite dependencies of each task. Our capability metrics (NC, IC, RC) score each task under unassisted, correct, or incorrect prerequisites to diagnose where failures arise; contribution metrics (N-Score, S-Score), adapted from probabilities of causation, quantify each prerequisite's necessity and sufficiency to determine why. We instantiate the framework in CADET, a diagnostic benchmark of 10 comp
קרא במקור המקורי