כתבה
arXiv cs.LG ·
למידה לזכור: פילוסופיה לשמירת זיכרון למודלי רשת חוזרת קצרים
Learning to Remember: Distilling Memory Retention for Compact Recurrent Neural Networks
מאמר חדש עוסק בפיתוח פרוטוקול לשמירת זיכרון במודלי רשת חוזרת קצרים. הפרוטוקול, המכונה MemKD, מאפשר קיצור קודם של המודל תוך שמירת יכולתו לזכור פרטים. המאמר כולל ניסויים וביצועים של MemKD.
תקציר מקורי באנגליתarXiv:2610.06942v1 Announce Type: new Abstract: Deep learning models, particularly recurrent neural networks and their variants, such as long short-term memory, have significantly advanced time series analysis. These models capture complex, sequential patterns in time series, enabling real-time assessments. However, their high computational complexity and large model sizes pose challenges for deployment in resource-constrained environments, such as wearable devices and edge computing platforms. Knowledge Distillation (KD) offers a solution by transferring knowledge from a large, complex model (teacher) to a smaller, more efficient model (student), thereby retaining high performance while reducing computational demands. Current KD methods, originally designed for computer vision tasks, negl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית