כתבה
arXiv cs.LG ·
Right In-Place (RiP) Convolution: סטרטגיה פשוטה ואופטימלית ל-CNN Inference
Right In-Place (RiP) Convolution: A Simple, General, and Near-Optimal Strategy for Memory-Efficient CNN Inference
סטרטגיה חדשה ל-CNN Inference המשפרת את כמות המספריים הפעילים. הסטרטגיה, הנקראת Right In-Place (RiP) Convolution, מאפשרת ל-CNN לבצע את ה-CNN Inference באופן יעיל יותר, על ידי שימוש במקום זמין.
תקציר מקורי באנגליתarXiv:2610.00586v1 Announce Type: new Abstract: Activation memory, not compute, limits CNN inference on constrained hardware such as microcontrollers. Direct in-place convolution removes the dual-buffer cost, but the memory-optimal formulation of Gural and Murmann assumes valid padding, unit stride, unit dilation, and odd square kernels, and needs a non-sequential traversal costing $2\times$ inference time in transposes. We identify two regimes in which their published closed form does not hold: (1) an under-allocation of exactly $(k-1)C_{in} \bmod (C_{out}-C_{in})$ scalars, active on every convolutional layer of their own deployed network and manifesting as a silent corruption of still-live input; (2) an unbounded overestimate, up to $2{,}432\times$, once the critical leg leaves the outpu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית