כתבה
arXiv cs.LG ·
התאמת מודלים להיגיון באמצעות הדרכה ומיזוג
Leveraging Instruction Tuning and Merging for Reasoning Model Adaptation
חוקרים מציגים שיטה להתאמת מודלים להיגיון באמצעות הדרכה ומיזוג. השיטה מאפשרת שימוש בנתונים מועטים וחסכון בעלויות. היא נבדקה על מודלים שונים והראתה שיפורים משמעותיים.
תקציר מקורי באנגליתarXiv:2607.14895v2 Announce Type: replace Abstract: Reasoning language models (RLMs) demonstrate impressive performance by leveraging test-time compute in the form of reasoning tokens. However, this behavior makes adapting RLMs to new domains challenging and expensive. The reason is that further training can disturb the learned behavior and degrade model performance. This makes it difficult to leverage supervised fine-tuning data with human-written solutions: although it contains high-quality annotations, it lacks reasoning tokens. In this work, we show how, despite this challenge, such data can be used efficiently for RLM adaptation. For this, we first use standard instruction tuning. Next, we leverage model merging to combine the instruction-tuned model with the original RLM, picking the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית