יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

LatentHarness: למידת אקציונים רציפים לזיכרון ולתפישה דרך תרגום-פוליצי

LatentHarness: Learning Latent Actions for Memory and Reasoning via Counterfactual Policy Distillation
LatentHarness מלמד אקציונים רציפים לזיכרון ולתפישה דרך תרגום-פוליצי. המודל חוקר את האפשרות לשיפור בעבודת-זיכרון ובעבודת-תפישה. LatentHarness נבחן על שישה מבחנים שונים והוכיח יתרונות ניכרים. המודל נמצא יעיל יותר ומהיר יותר מהבסיסים החזקים.
תקציר מקורי באנגליתarXiv:2609.39740v1 Announce Type: new Abstract: Long-context reasoning faces two complementary bottlenecks: retaining evidence across long inputs and sustaining computation across many reasoning steps. Existing approaches largely address them separately, with external memory extending access to distant evidence and latent reasoning compressing multi-step computation. We introduce LatentHarness, which unifies memory access and latent reasoning as sequential latent action selection. At each internal step, the model chooses THINK for further computation, RECALL from a fast-weight memory of input evidence and intermediate reasoning states, or EXIT to emit the next token. We train this policy with counterfactual policy distillation, which branches every action for one step and scores its effect
קרא במקור המקורי