כתבה
arXiv cs.LG ·
Forking: Sudden Overfitting Under Replay
תקציר מקורי באנגליתarXiv:2610.00394v1 Announce Type: new Abstract: This paper studies forking, a generalization failure discovered in NanoGPT autoresearch. Under data replay, models with an over-encoding n-gram memory branch show a sharp separation of training and validation loss at epoch boundaries, resembling the shape of forks. We study this phenomenon in a controlled vanilla NanoGPT setting and reproduce it in a DeepSeek-style model with Engram. Mechanistically, repeated updates sharpen the continuations observed in training while suppressing the probability of unseen continuations, whose loss grows with each pass. The n-gram module creates weakly interacting context-specific subspaces, amplifying this effect. Low-frequency contexts contribute most of the gap, whereas larger training budgets and heavily
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית