יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

CurveTQ: קיצור טרליס ללא הסתברות של כונני LLM דרך חיפוש עם כונני משקל

CurveTQ: Rotation-Free Trellis Quantization of LLM Weights via Curvature-Weighted Search
מחקר חדש: CurveTQ, טכניקה לקיצור טרליס של כונני LLM, משפרת את דיוק המודלים ב-1-3 נקודות. הטכניקה, שפותחה על ידי צוות מדעניות ומדעני טכנולוגיה, נועדה לשפר את יעילות המודלים ולאפשר להם להתאים למערכות קצרות-זיכרון.
תקציר מקורי באנגליתarXiv:2610.09212v1 Announce Type: new Abstract: The best two-bit weight quantizers for large language models, such as QTIP and Proteus, rotate each weight matrix by a random orthogonal transform, which must be undone at every decoding step, then encode it with a trellis or lattice code under a Euclidean search; the layer Hessian enters only through error feedback between coding blocks. We show that this leaves part of the Hessian unused. Error feedback turns the loss into a weighted sum of per-coordinate rounding errors whose weights, the diagonal of the Hessian's LDL factorization, existing quantizers compute but never read. We put these weights into the Viterbi branch metric, so the search follows the curvature within each coding block. This also explains the rotation: it removes this wi
קרא במקור המקורי