יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

XOR-Trellis: קוונטיזציה חדשה למודלי LLM

XOR-Trellis: Ultra-Low-Complexity Dequantization and Curvature-Aware Hadamard-Free LLM Quantization
XOR-Trellis הוא אלגוריתם חדש לקוונטיזציה של מודלי LLM. הוא מאפשר דחיסה גבוהה של משקלי המודל ברוחב סיביות נמוך מאוד, ללא צורך בקודבוקים גדולים. האלגוריתם משתמש בטכניקות חדשות כדי לשפר את הדיוק ואת היעילות.
תקציר מקורי באנגליתarXiv:2610.00432v1 Announce Type: cross Abstract: Trellis-coded quantization enables high-dimensional compression of large language model (LLM) weights at ultra-low bit widths without the exponentially large codebooks required by conventional vector quantization. Practical deployment, however, presents two challenges: reconstructing compressed weights at sufficient parallel throughput to avoid making dequantization an inference bottleneck, and maintaining quantization accuracy without costly incoherence transformations. We address these challenges with two complementary techniques. First, we introduce an ultra-low-complexity trellis dequantizer that uses a structured, hardware-efficient state-to-value mapping while preserving diverse reconstruction choices for trellis search. Second, we re
קרא במקור המקורי