יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיפור של דגלי שפה גדולים באמצעות תלת-קווי תלוי-מסכה

Boosting Large Language Models with Mask Fine-Tuning
במאמר זה, המחברים מציגים פרדיגמה חדשה להטמעת דגלי שפה גדולים, המכונה Mask Fine-Tuning (MFT). MFT לומד ומיישם מסכות בינאריות על דגלי שפה גדולים, תוך שימוש במטרייה הסטנדרטית של LLM. המחברים מציגים תוצאות של MFT על דגלי שפה גדולים שונים, כולל LLaMA2-7B ו-Llama3.1-8B.
תקציר מקורי באנגליתarXiv:2503.22764v3 Announce Type: replace-cross Abstract: The large language model (LLM) is typically integrated into the mainstream optimization protocol. However, it remains underexplored whether maintaining the model integrity is \textit{indispensable} for promising performance. In this work, we introduce Mask Fine-Tuning (MFT), a novel LLM fine-tuning paradigm demonstrating that carefully breaking the model's structural integrity can surprisingly improve performance without updating model weights. MFT learns and applies binary masks to well-optimized models, using the standard LLM fine-tuning objective as supervision. Based on fully fine-tuned models, MFT uses the same fine-tuning datasets to achieve consistent performance gains across domains and backbones (e.g., an average gain of 2.
קרא במקור המקורי