כתבה
arXiv cs.AI ·
When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue
תקציר מקורי באנגליתarXiv:2608.27176v2 Announce Type: replace-cross Abstract: Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. We formalize this failure mode as cross-modal disagreement, where transcripts suggest plausible but incorrect surface interpretations while acoustic cues such as prosody or speaking style support different answers. We develop a scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples. We also include consistent cases where transcript-based and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית