יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

פרוזודיה לטקסט: ניבוי טקסט מדיבור מסונן

Prosody-to-Text: Predicting text from low-pass filtered speech
חוקרים פיתחו שיטה לניבוי טקסט מדיבור מסונן. השיטה משתמשת במודל Whisper ומסננת את 12 הבינים הנמוכים ביותר. התוצאות היו מדויקות, עם 10% מההגעות שהוחזרו במלואן ו-40% עם שגיאת מילים ממוצעת של 25% או פחות.
תקציר מקורי באנגליתarXiv:2610.11544v1 Announce Type: new Abstract: While predicting prosody from text is an established task in the field, the opposite direction, predicting text that fits a given prosodic pattern, remains largely overlooked. We find this unfortunate, because this opposite direction could lead to some very interesting use cases. Therefore, in this paper, we make the first steps in the prosody-to-text direction by inves- tigating how much of the original sentence can be recovered from its prosodic pattern. To this end, we fine-tune the Whis- per model using only the 12 lowest Mel bins (low-pass filter with approximately 450Hz cutoff), and obtain surprisingly accurate results (WER 36%), with 10% of utterances be- ing recovered perfectly, and 40% of utterances having Word Error Rate at or below
קרא במקור המקורי