כתבה
arXiv cs.AI ·
פער תוכן רציף בייצוגי טקסט
Isotropic Yet Undecodable: The Sequential Content-Sufficiency Gap in Latent-Predictive Text Representations
חוקרים את הפער בתוכן רציף בייצוגי טקסט. הם מציגים מסגרת חדשה, CANOPE, שמשלבת קנבסים רציפים והדרכה גאומטרית. התוצאות מראות שקנבסים אלו משפרים את הדיוק בטקסטים מושחתים.
תקציר מקורי באנגליתarXiv:2610.07906v1 Announce Type: new Abstract: We study sequential content sufficiency by investigating whether a representation retains the ordered target information available in its input. An information-theoretic decomposition separates input ambiguity, representation loss, and readout mismatch. We construct recoverable views where perfect agreement and joint isotropic Gaussianity coexist with zero target information, and establish limits imposed by deterministic canonical anchors. Token log-loss provides a one-sided information-loss bound; a fixed-penalty ridge analysis shows why rank alone cannot determine prediction risk. These results motivate CANOPE, a nonautoregressive framework with ordered latent canvases, canonical-token supervision, and geometric regularization. On 40,000 va
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית