כתבה
arXiv cs.CL ·
LIFT: מודל שפה עם משוב עמוק
Pretraining Latent Information Feedback Transformers with Teacher Supervision
LIFT הוא מודל שפה חדש שמאפשר משוב עמוק בין שכבות. המודל מאומן עם הדרכה של מורה, ומצליח לשפר ביצועים במגוון משימות. LIFT משתמש במודלים קיימים כמו LLaMA.
תקציר מקורי באנגליתarXiv:2609.38149v1 Announce Type: new Abstract: Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-she
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית