כתבה
arXiv cs.AI ·
Trajectory-Retrieval Speculative Decoding: כאשר ההיסטוריה של המודל עוזרת?
Trajectory-Retrieval Speculative Decoding: When Does a Model's Own History Help?
אנו מחקרים את הרגעים בהם ההיסטוריה של המודל עוזרת בתהליך הפיענוח. המאמר עוסק בשימוש בהיסטוריה של המודל כדי לשפר את תהליך הפיענוח.
תקציר מקורי באנגליתarXiv:2610.07350v1 Announce Type: new Abstract: Long chain-of-thought reasoning increases sequential decoding cost while creating a growing history of potentially reusable continuations. We investigate when this history supplies useful drafts and complements an existing drafter. Controlled source comparisons reveal trajectory-specific reuse, motivating our method Trajectory-Local Adaptive Retrieval (TLAR). TLAR retrieves approximately matched continuations from the current trajectory and uses recent verification outcomes to adapt retrieval activation and candidate width. TLAR combines retrieved continuations with model-generated drafts in a shared candidate tree, preserving the target model's output distribution through exact verification. Across code debugging, mathematics, and open-ended
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית