יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שמעו את הלטנטים: הכרזה עצמית של הזיהוי הדיבורי בדגימות קוליות גדולות באמצעות התקשרות במצבי החבויים

Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
במאמר זה, נציגים טכניקה חדשה לשיפור זיהוי דיבור בדגימות קוליות גדולות, על ידי שימוש בהתקשרות במצבי החבויים. הטכניקה, הקרויה Hybrid Search, משתמשת במודלי גמיני ו-LangGraph לשיפור זיהוי הדיבור.
תקציר מקורי באנגליתarXiv:2609.02940v1 Announce Type: new Abstract: Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two strategies remains underexplored. In this work, we refine warm-initialized LLM-based ASR models by leveraging their own pre-adaptation base LLMs, focusing on LoRA-adapted settings where the base LLM is preserved. To achieve this, we propose Hybrid Search, a targeted correction strategy motivated by two observations. First, interaction features that characterize the relationship between LLM-based ASR hidden states and base-LLM hidden states provide informative signals about a toke
קרא במקור המקורי