כתבה
arXiv cs.LG ·
Backward-State Policy Is Part of the Learning Algorithm
תקציר מקורי באנגליתarXiv:2609.39813v1 Announce Type: new Abstract: Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, the original, or a new random rounding. This backward-state policy looks like a memory and precision detail, settled by copy accuracy and final loss. We argue that it is part of the learning algorithm, and that neither check shows whether it is right. Copy accuracy does not decide the outcome: in three pairs of 390M runs with an emulated FP8 backward, training fails when attention's backward reuses the forward's rounded output and succeeds with a new rounding from the same distribution. Even the most accurate copy, the original itself, can be wrong by our reference: the gradient of the forward p
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית