כתבה
arXiv cs.LG ·
Decision Titan: טירוף בזמן המבחן לזיכרון ארוך-טווח בלמידה עצמית רפלקסיבית
Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
במאמר זה, נחקר האפשרות של טירוף בזמן המבחן (Test-Time Training) לזיכרון ארוך-טווח בלמידה עצמית רפלקסיבית. המאמר עוסק בפיתוח דגם חדש, Decision Titan, המשלב טירוף בזמן המבחן עם דגם Decision Transformer.
תקציר מקורי באנגליתarXiv:2610.01513v1 Announce Type: cross Abstract: Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has proposed tackling this problem with the Test-Time Training (TTT) framework, which stores episodic memories in the parameters of a neural network through gradient descent at both train and test-time. This approach has seen success in the domain of Natural Language Processing, however, to the best of our knowledge it has not yet been applied to the domain of Reinforcement Learning (RL), nor has there been a study analysing how this memory practically fu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית