יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שרדינג עדין לפי קרבה לפי תאריך

Affinity-Aware Sharding for Delayed Tensor Parallelism
DTP (Delayed Tensor Parallelism) משדרג את שרדינג המודלים של Transformer. המחקר עוסק בשרדינג עדין של מודלי Transformer, ומציע שיטה לשרדינג עדין שמקסימה את קרבת הקבוצות.
תקציר מקורי באנגליתarXiv:2609.13846v1 Announce Type: new Abstract: Delayed Tensor Parallelism (DTP) removes the blocking all-reduce of tensor-parallel Transformer inference. Every device adds its own partial output to its residual stream (and broadcasts it) immediately, but only gathers (receives) the other devices' partials $\delta$ modules later. A TP to DTP change therefore amounts to a real architecture change, and dense Transformer models need to be retrained or distilled after adaptation. We show that DTP breaks the permutation symmetry of neurons inside FFNs and of KV heads inside attention modules, and that this symmetry breakage makes the sharding itself a modelling decision. We show that maximising the affinity between the KV heads and the FFN neurons co-located on a device, by permuting the dense
קרא במקור המקורי