כתבה
arXiv cs.CL ·
DirectSpeech2LLM: פלטפורמה פשוטה למניעת עדכנות פרסומים ב-LLM
DirectSpeech2LLM: A Simple End-to-End Framework to Mitigate Prompt Overfitting in Speech-LLMs
פלטפורמה שמטפלת בעדכנות פרסומים ב-LLM לשפה, ומאפשרת התאמה לתרגום ולהבנת רגשות.
תקציר מקורי באנגליתarXiv:2610.08085v1 Announce Type: new Abstract: Speech-LLMs often exhibit prompt overfitting, where models solely trained on automatic speech recognition (ASR) instruction fail to generalize to new instructions such as speech translation and continue to behave primarily as ASR system. We propose DirectSpeech2LLM, a simple end-to-end framework that preserves the instruction-following ability of the LLM on unseen tasks when conditioned on speech. It computes distance-based CTC loss over the frozen LLM embedding matrix and uses greedy CTC labels to derive geometrically and temporally aligned speech embeddings respectively as an input to the LLM. Trained solely on 960 hours of LibriSpeech ASR data, DirectSpeech2LLM outperforms the cascaded system on ASR (seen task) and generalizes zero-shot to
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית