כתבה
arXiv cs.CL ·
EchoDistill: רובוסטיות של דגמי שפה גדולי-קול תחת עיוות-לנק-עצמי
EchoDistill: Robust Large Audio Language Models via Noisy-to-Clean Self-Distillation
דגמי שפה גדולי-קול רובוסטיים תחת עיוות-לנק-עצמי. EchoDistill משתמש באודיו טהור כמידע פריבילגיות בעת הכשרה-אחרי-למידה.
תקציר מקורי באנגליתarXiv:2605.23954v2 Announce Type: replace Abstract: Large Audio Language Models (LALMs) remain vulnerable to acoustic noise, which can obscure task-relevant evidence and produce unreliable responses. We propose EchoDistill, a noisy-to-clean self-distillation framework that uses clean audio as privileged information during post-training. A noisy-input student samples candidate responses reflecting its inference-time behavior, while a frozen copy of the same backbone processes the corresponding clean audio. EchoDistill combines masked response-token distillation, task-gated consistency shaping, and teacher-referenced group-relative optimization to align noisy-input generation with clean-conditioned semantics. Only the student is retained at inference time, introducing no additional inference
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית