יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מערכת ASR דו-צורה: תפיסה-מודע ASR להפיכה של טקסט כתוב לטקסט דיבור

Dual-Form ASR: Semantics-Aware Inverse Text Normalization for Chinese Speech Recognition
מערכת ASR דו-צורה שמסוגלת להפיך טקסט כתוב לטקסט דיבור, תוך שימוש בטכנולוגיית LLM ובאלגוריתם ITN-MWER.
תקציר מקורי באנגליתarXiv:2609.02901v1 Announce Type: new Abstract: Modern automatic speech recognition (ASR) scenarios require both spoken-form transcripts for faithful transcription and readable written-form transcripts with inverse text normalization (ITN). However, these forms are typically produced by cascaded modules, where a spoken-form ASR output is rewritten by a separate ITN component, making written-form ASR-ITN vulnerable to recognition errors and decoupling normalization from acoustic-contextual modeling, especially for semantically dependent numeric expressions. In this paper, we propose Dual-Form ASR (DF-ASR), a framework that extends spoken-form ASR capability to semantics-aware written-form ITN through paired spoken-form and written-form supervision while retaining prompt-level selection betw
קרא במקור המקורי