כתבה
arXiv cs.LG ·
Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware
תקציר מקורי באנגליתarXiv:2610.00687v1 Announce Type: cross Abstract: Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training. We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית