כתבה
arXiv cs.CL ·
GIVE-KWS: גישה מבוקרת לאביזרי ראייה ויזואליים לזיהוי קלט-מסגרת של מילות דיבור
GIVE-KWS: Gated Injection of Visual Evidence for Noise-Robust Query-by-Example Keyword Spotting
GIVE-KWS מציע זיהוי קלט-מסגרת של מילות דיבור עם רובוסטנסיית רעש, על ידי הזרמת אביזרי ראייה ויזואליים. המערכת, GIVE-KWS, נבחנה בבסיס ניסויי QbyE-KWS, והוכיחה יעילות גבוהה בהשוואה למערכת טקסט-ו-אודיו. GIVE-KWS נבנה על ידי גישה מבוקרת לאביזרי ראייה ויזואליים, GIVE (Gated Injection of Visual Evidence), שמציעה זיהוי קלט-מסגרת של מילות דיבור עם רובוסטנסיית רעש.
תקציר מקורי באנגליתarXiv:2610.07046v1 Announce Type: cross Abstract: Visual speech promises noise-robust keyword spotting, yet a visual stream is not necessarily used. On a tri-modal query-by-example keyword spotting (QbyE-KWS) benchmark, we find that a system with a task-trained visual encoder comes within 2 percentage points of a text-and-audio system in equal error rate (EER) at -10 dB, and link this gap to the encoder's lack of phonemic information. We present GIVE-KWS, whose fusion stage, GIVE (Gated Injection of Visual Evidence), conditions query audio on lip motion through gated cross-attention. We show that visual robustness depends on two interacting conditions: a phoneme-bearing visual representation, and fusion that injects visual evidence rather than rescaling audio features. Under a phoneme-bear
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית