יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שימור מיניבאטץ' - 8 שנים לאחר מכן: מה עלות שימור ומה היא יוצרת בנתונים

Minibatch persistency, eight years later: what batch reuse costs in steps and joules, and what it saves in data
שימור מיניבאטץ' - זאת העלות והיתרון של שיטה זו בנתונים. המאמר עורך ניסויים ומציג תוצאות.
תקציר מקורי באנגליתarXiv:2609.13922v1 Announce Type: new Abstract: Minibatch persistency reuses data instead of reading it: rather than drawing a fresh minibatch at every optimizer step, it takes K consecutive steps on the same one. Absorbed into data echoing in 2019, it has carried one objection -- that reuse merely imitates a larger learning rate -- and no baseline tuned as carefully as the method itself. This paper runs the missing test. A pre-registered study trains a 49M-parameter Transformer on FineWeb-Edu at minibatch size B in {32, 128, 512}, 8 seeds per cell, tuning the learning rate separately for every batch size and every arm, against a reuse-free control that changes the sampling and nothing else. Each headline claim is a cost to reach a fixed loss, read on four axes: optimizer steps, fresh toke
קרא במקור המקורי