כתבה
arXiv cs.AI ·
קומפרציה של כוח: כמה ניצבות דטרמיניזציה ולא כמה כיוון של כיוון?
Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression
קומפרציה של כוח: כמה ניצבות דטרמיניזציה ולא כמה כיוון של כיוון? נמצא כי קומפרציה של כוח גורמת לדטרמיניזציה, ולא לכיוון של כיוון.
תקציר מקורי באנגליתarXiv:2609.07901v1 Announce Type: new Abstract: Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4-bit quantization of mid-sized models is often considered "nearly free." We examine a different question: when several answers are valid, does quantization change what a model chooses to say? We serve three checkpoints (Qwen3-8B/14B/32B) at three weight precisions (W4A16 AWQ, W8A16 FP8-Marlin, and bf16), holding the hardware, software, and sampling configuration constant, and collect approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries. We pre-specified the analyses in three waves in version control. At 8B, int4 reduces output diversity: th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית