כתבה
arXiv cs.CL ·
דגימות דיבור עצמאיות גילו אריתמטיקה של וקטורי פונולוגיה
[b] = [d] - [t] + [p]: Self-supervised Speech Models Discover Phonological Vector Arithmetic
דגימות דיבור עצמאיות גילו שהן משתמשות באריתמטיקה של וקטורי פונולוגיה. כל הקוד והדמואים האינטראקטיביים זמינים ב-https://github.com/juice500ml/phonetic-arithmetic.
תקציר מקורי באנגליתarXiv:2602.18899v4 Announce Type: replace-cross Abstract: Self-supervised speech models (S3Ms) are known to encode rich phonetic information, yet how this information is structured remains underexplored. We conduct a comprehensive study across 96 languages to analyze the underlying structure of S3M representations, with particular attention to phonological vectors. We first show that there exist linear directions within the model's representation space that correspond to phonological features. We further demonstrate that the scale of these phonological vectors correlate to the degree of acoustic realization of their corresponding phonological features in a continuous manner. For example, the difference between [d] and [t] yields a voicing vector: adding this vector to [p] produces [b], whi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית