כתבה
arXiv cs.LG ·
The Conflict Between Logic and Memory: Training Conditions for Optimizer-Dependent Rule Acquisition
תקציר מקורי באנגליתarXiv:2610.00403v2 Announce Type: replace Abstract: Optimizers can fit the same task while acquiring different generalizing relations. We study the training conditions governing these differences in single-hidden-layer ReLU networks, combining composite evidence tasks, parameter-level interventions, and a three-seed strict-parity scan. Our central finding is that nuisance-connected trainability reshapes both shared failures and relative optimizer advantages. In a nuisance-heavy task, all twenty tested optimizer configurations remain near chance on the hardest stage. Retaining every input but fixing nuisance-connected first-layer weights at initialization raises that stage's accuracy from approximately 50\% to 70.56\%, 68.47\%, and 69.00\% for momentum SGD, Adam, and Muon. Masking the same
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית