יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

APEX: תהייה חכמה, לא עמוקה

APEX: Speculate smarter, not deeper
APEX הוא בקר למודל שמאפשר מהירות גבוהה יותר בזמן אינפרנס. הוא משתמש במנגנון תהייה והתאמה של עומק הניחוש. APEX מוכלל במודל vLLM ונבדק עם Qwen3-8B, והשיג עד 5.24X מהירות יותר מאשר פענוח אוטורגרסיבי.
תקציר מקורי באנגליתarXiv:2610.07780v1 Announce Type: new Abstract: Speculative decoding reduces large language model inference latency by drafting multiple tokens before target-model verification, but its effectiveness depends on both the proposal mechanism and draft depth. Fixed configurations cannot respond to changes in predictability, repetition, and acceptance during generation, so deeper drafting can increase wasted computation without proportional speedup. We introduce APEX, a learned controller that balances decoding speed and draft-token waste through request-level expert selection and block-level depth adaptation. APEX-Router selects among EAGLE-3, n-gram, and draft-model speculation for each request, while APEX-Depth adjusts draft length at each verification block using causal decoding signals and
קרא במקור המקורי