כתבה
arXiv cs.AI ·
EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
תקציר מקורי באנגליתarXiv:2605.23954v2 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the corresponding clean audio. EchoDistill combines masked response-token distillation, task-gated consistency shaping, and teacher-referenced group-relative optimization to align noisy-input generation with clean-conditioned semantics. Only the student is retained at inference time, introducing no additional inf
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית