יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

APEX: Speculate smarter, not deeper

תקציר מקורי באנגליתarXiv:2610.07780v1 Announce Type: cross Abstract: Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals a
קרא במקור המקורי