יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

האם קישור כבדות משקל עדיין יעיל למודלי LLM רק-מקבלי-קלט בסביבה פרטית תחת DP-SGD?

Is Weight Tying Still Beneficial for Decoder-Only LLMs in Private Settings Under DP-SGD?
במאמר זה נבחן השפעת קישור כבדות משקל על יעילות מודלי LLM רק-מקבלי-קלט בסביבה פרטית תחת DP-SGD. התוצאות הראו שאין צורך בקישור כבדות משקל ושאפשרות זו יכולה להפחית את צריכת המשאבים.
תקציר מקורי באנגליתarXiv:2609.40335v1 Announce Type: new Abstract: Differentially Private Stochastic Gradient Descent (DP-SGD) is a leading approach for privacy-preserving fine-tuning of large language models (LLMs). Many decoder-only LLMs employ weight tying between input and output embeddings, a design choice originally introduced for parameter efficiency and improved language modeling performance in the non-private setting. However, the impact of weight tying under differentially private training remains largely unexplored. In this work, we investigate the role of weight tying in the DP setting using GPT2 and DistilGPT2 as representative decoder-only architectures. Interestingly, we find that untied embeddings consistently outperform weight-tied models under DP-SGD, achieving gains of up to 4.74% points i
קרא במקור המקורי