יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

DeltaTTT: אופטימיזציה שכבתית לזיכרון רקורנטי לא ליניארי

DeltaTTT: Layerwise Optimization for Nonlinear Recurrent Memory
DeltaTTT היא שיטה חדשה לאופטימיזציה של זיכרון רקורנטי לא ליניארי. היא מאפשרת חישוב מקבילי ושיפור בביצועים. ניסויים על DeltaNet ו-LaCT הראו שיפורים בדגמי שפה ואחזור.
תקציר מקורי באנגליתarXiv:2610.08553v1 Announce Type: cross Abstract: Sequential test-time training adapts a memory network through successive updates, each computing an inner-loop gradient based on the network's previous state. Intuitively, this state dependence should allow each update to account for what the memory has already learned and better incorporate new information. However, we find that this expected advantage does not consistently materialize in nonlinear memories: a fixed-base parallel TTT baseline outperforms its serial counterpart. Our exploratory experiments point to a key underlying difficulty: nonlinear memories can be harder to optimize than linear ones within a single pass over the sequence. To alleviate this optimization difficulty, we introduce DeltaTTT, which replaces joint inner-loop
קרא במקור המקורי