כתבה
arXiv cs.LG ·
הגרדיאנט אינו רואה דירוג: חוסר תלות בדירוג במטריצת-CODI
The Gradient Does Not See Rank: Rank-Indifference in Matrix-CODI on ProsQA
מודל Matrix-CODI הראה חוסר תלות בדירוג. המחקר בדק את השפעת הדירוג על ביצועי המודל ומצא שאין הבדל משמעותי. התוצאות מראות שהמודל אינו תלוי בדירוג המטריצה.
תקציר מקורי באנגליתarXiv:2609.03090v1 Announce Type: new Abstract: Continuous chain-of-thought models compress reasoning into latent tokens. Matrix-valued variants, which route each latent token through a d x d matrix bottleneck, introduce rank as a single-sample structural observable on the latent matrix Z. If matrix latents carry parallel reasoning paths via superposition, rank should track them, and truncating Z to low rank should hurt accuracy on tasks whose solutions plausibly require multiple components. Across four training regimes of a matrix-CODI model (three on ProsQA, one on GSM8K-Aug below the learning threshold), the rank-k projection ablation curve is flat to within 0.6 percentage points. A three-seed replication yields 81.0 +/- 2.0 percentage points accuracy while the final effective rank of Z
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית