כתבה
arXiv cs.AI ·
What Does a Token Cost? A Mixture-of-Agents Measurement of Sufficient Per-Token Compute
תקציר מקורי באנגליתarXiv:2610.02491v1 Announce Type: new Abstract: Large language models spend the same amount of computation on every token they generate, regardless of how difficult each token is to produce. Methods such as speculative decoding and model routing are built on the premise that much of this computation is unnecessary, yet the computation an individual token actually requires has not been measured. We measure it through a Mixture-of-Agents (MoA) lens: a panel of fifteen language models of increasing capacity, drawn from three families, in which every agent attempts to reproduce a reference sequence token by token, conditioned on the correct preceding tokens. We define the inference cost of the smallest agent that succeeds as the token's sufficient compute, which upper-bounds what the token req
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית