כתבה
arXiv cs.AI ·
X-AuT: קיצור פרוגרסיבי של קודקודד-אודיו ל-LLMs עם צינון-סקאלה
X-AuT: Progressive Audio-Encoder Compression for Speech LLMs with Cross-Scale Distillation
X-AuT היא פרוגרסיה של קיצור פרוגרסיבי של קודקודד-אודיו ל-LLMs עם צינון-סקאלה. היא מקצרת את העומס של ה-LLMs ומשפרת את הדיוק שלהם. X-AuT נבחנה על 10 מבחנים ציבוריים והיא הציגה תוצאות טובות. הפרויקט נמצא באתר https://xpeng-ai.github.io/x-aut.
תקציר מקורי באנגליתarXiv:2609.11412v1 Announce Type: cross Abstract: Reducing audio-encoder depth lowers the inference cost of speech large language models, but removing complete blocks perturbs the embeddings consumed by the decoder and can cause deletion and premature end-of-sequence errors. We introduce X-AuT, a progressive framework that selects layer combinations through short behavioral probes and restores the pruned model through representation alignment, cross-scale distillation, scheduled student-policy supervision, and LoRA finetuning. The language-model backbone remains frozen, while attention LoRA adapters and the tied output embedding adapt during distillation. Training uses the highest-agreement tier from a transcript-consistency pipeline, followed by source reweighting during finetuning. On te
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית