כתבה
arXiv cs.LG ·
רונדינג סטוכסטי באינפרנסה של תרגומן תחת-דינמי: מחקר על חקירה של GPT-2
Stochastic Rounding in Low-Precision Transformer Inference: A Variable-Precision Emulation Study of a Small GPT-2
מחקר על רונדינג סטוכסטי באינפרנסה של תרגומן תחת-דינמי ל-GPT-2. המחקר חקר את השפעת רונדינג סטוכסטי ורונדינג לקרוב ביותר על תרגומן תחת-דינמי של GPT-2. התוצאות הראו שרונדינג סטוכסטי יכול להפחית את הפליקסיטי ל-1.10x מהתרגומן המלא-דינמי.
תקציר מקורי באנגליתarXiv:2610.01889v1 Announce Type: new Abstract: Should low-precision transformer inference use stochastic rounding (SR) or round-to-nearest (RN)? The answer depends on where in the network you look. We isolate this effect by holding the numerical format fixed and varying only the rounding rule at individual operation sites. To enable experiments at freely chosen precisions, we extend the PRISM vectorized rounding library to arbitrary virtual precision via a variable-precision stochastic rounding (VPSR) algorithm, proving that the rounding decision is evaluated exactly in hardware floating point. We develop two analyses providing complementary insight into this site-level trade-off. First, a probabilistic forward-error bound for linear projections shows that SR's error envelope grows as $O(
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית