כתבה
arXiv cs.CL ·
שדים בשאילת העברה
Devils in Question Relay: Source-Conditioned Relay Steering to Mitigate Hallucinations in Audio-visual Large Language Models
מחקר חדש מציג שיטה למניעת הזיות במודלים של שפה גדולים רב-מודאליים. השיטה, SECRET, משתמשת בניתוח התיווך והצגה של המודל כדי לזהות ולתקן הזיות. הניסויים הראו שיפור משמעותי בביצועים.
תקציר מקורי באנגליתarXiv:2609.37568v1 Announce Type: new Abstract: Audio-visual large language models (AVLLMs) have made remarkable progress in multimodal understanding and reasoning through interactions among visual, auditory, and linguistic information. However, recent studies show that AVLLMs face a critical challenge: $\textbf{source-confused grounding hallucination}$, where cues from the unused modality induce responses that the required modality does not support, undermining reliability in real-world applications. Existing methods have made progress in mitigating this failure, yet how it arises from internal cross-modal interactions remains insufficiently understood. To address this gap, we conduct path-intervention and representation analyses, revealing a $\textbf{question-relay}$ mechanism: question
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית