כתבה
arXiv cs.LG ·
Sharp Rates and a One-Line Correction for Spectral Representation Learning
תקציר מקורי באנגליתarXiv:2609.15825v1 Announce Type: new Abstract: A self-supervised encoder is trained once, frozen, and reused through lightweight probes on tasks nobody named at training time; the practitioner's question is when the off-the-shelf features are good enough and when they need fixing. Canonical correlation analysis, HGR maximal correlation, and the population optimum of the spectral contrastive loss all return the top-$k$ singular subspace of a cross-view dependence operator, justified by isotropy: if the task prior has no directional preference, that subspace is universally optimal. We show isotropy is the wrong hypothesis. The prior enters the transfer risk only through the task covariance $\Lambda=\mathbb{E}[\Delta\Delta^\top]$, and only through its compression onto the operator's leading
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית