כתבה
arXiv cs.AI ·
מודלי LLM אודיו יודעים מתי הם לא שומעים
Audio LLMs Know When They Can't Hear You
חוקרים פיתחו שיטה לזיהוי תעתיקים לא אמינים במודלי LLM אודיו. השיטה מאפשרת למודלים לזהות מתי הם לא שומעים נכון ולבקש שיפור. המחקר משתמש במודל Llama.
תקציר מקורי באנגליתarXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, we study model-conditional transcription reliability: whether an Audio LLM can recognize when its own transcription is unreliable. We first prompt the Audio LLM to assess whether its own transcription would be reliable, and find that the model is a poor judge of its own transcription reliability: in most cases, it predicts that its transcription will be reliable. We find that existing approaches, including speech quality predictors, audio LLM generation uncertainty, and transcript-conditioned WER estimation, provid
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית