יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

צמצום תעתיקים מדומים ב-Whisper

Reducing Hallucinated Transcripts in Whisper via Hallucination Space Projection
חוקרים הצליחו לצמצם תעתיקים מדומים במודל Whisper לרכישת דיבור, באמצעות שיטה חדשה של פרויקציה. השיטה מורידה את שיעור התעתיקים המדומים ב-92.21%. המודל Whisper הוא מודל יסודי נפוץ לרכישת דיבור אוטומטית.
תקציר מקורי באנגליתarXiv:2609.04561v1 Announce Type: new Abstract: Whisper is a widely used foundation model for automatic speech recognition (ASR), but its generative decoder can produce fluent hallucinated transcripts for inputs containing little or no speech. We propose a training-free, inference-time method to reduce these hallucinations using low-rank projection of decoder activations. A compact hallucination-associated subspace is estimated from non-speech calibration data, and decoder hidden states are projected away from this subspace during inference. We evaluate two variants: always-on, which applies projection to all inputs, and gated, which applies it only when Whisper predicts that an input is likely non-speech. Across non-speech benchmarks, always-on projection reduces average hallucination rat
קרא במקור המקורי