כתבה
arXiv cs.AI ·
Accelerating Constrained Decoding with Token Space Compression
תקציר מקורי באנגליתarXiv:2605.29986v2 Announce Type: replace Abstract: To guarantee that an LLM's outputs conform to a specified structure, context-free grammar (CFG) decoding engines force the selection of next tokens to produce strings that conform to a given CFG. Current CFG-constrained decoding engines are highly optimized, but still suffer from the inherent costs arising from their massive per-step search space---i.e. the entire token vocabulary. This results in intractably high overhead for more complex CFGs, which is precisely the situation where CFG engines are most useful. In this paper, we introduce CFGzip, an offline technique for compressing the token search space, which massively reduces CFG engine overhead. In experiments, we report latency reduction of up to 75x during batched inference, cutti
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית