יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פיתוח תואם-גבולות ל-Sinkhorn Attention: גרדיאנטים של גשר דבש-אופקי

Block-Wise Differentiable Sinkhorn Attention: Tail-Refinement Gradients with a Gap-Aware Dustbin Bridge
פיתוח גרדיאנטים של גשר דבש-אופקי ל-Sinkhorn Attention על ציוד TPU. המחקר עוסק באופטימיזציה של תואם-גבולות ל-Sinkhorn Attention על ציוד TPU. התוצאות המוצגות הן גרדיאנטים של גשר דבש-אופקי, שהם גרדיאנטים של גשר דבש-אופקי שהופעלו על ציוד TPU.
תקציר מקורי באנגליתarXiv:2605.08123v3 Announce Type: replace Abstract: We study long-context balanced entropic optimal transport (OT) attention on TPU hardware through a stopped-base, fixed-depth tail-refinement surrogate. After a stopped $T$-step Sinkhorn solve, we unroll a short refinement tail and differentiate that surrogate exactly. For the reported $R=2$ TPU path, the backward pass contains four staircase plan factors. We prove an exact one-reference-tile schedule: the $R=2$ score cotangent is a single reference plan tile times an explicit modifier field built from vector cotangents and dual differences. This yields block-wise cost $O((T+R)LW)$, $O(Ld)$ input storage, and $O(L)$ additional HBM usage for fixed head dimension $d$ and band width $W$ on the balanced fixed-support path. We also formalize th
קרא במקור המקורי