כתבה
arXiv cs.LG ·
כל קבוצה היא קבוצת התקנה שלה: חיפוש גרדיאנט על-זמן-אמצעי לבחירת נתונים באימון מחדש של LLM
Every Batch Is Its Own Validation Set: Leave-One-Out Gradient Matching for Online Data Selection in LLM Fine-Tuning
במאמר זה, המחברים מציגים טכניקה חדשה לבחירת נתונים באימון מחדש של LLM, המשתמשת בחיפוש גרדיאנט על-זמן-אמצעי. הטכניקה, הקרויה LOOM, משתמשת בגרף המטריצה של הגרדיאנטים כדי לבחור את הדגמים המועדפים לאימון. המחברים מציגים תוצאות של LOOM על-אורך שלושה תרגילי אימון, ומראים ש-LOOM משפר את התוצאות של האימון בהשוואה לאימון על-כל הנתונים.
תקציר מקורי באנגליתarXiv:2610.00436v1 Announce Type: new Abstract: Online batch selection fine-tunes a language model on the most useful part of each candidate batch. Selectors that match the gradient of the candidate batch are attractive because they need no held-out data, yet they rarely beat training on the whole batch. We show why. In-sample gradient matching uses every example as part of its own target, so its objective credits each example with its own gradient noise. This is the covariance penalty that makes training error optimistic, now sitting on the diagonal of the gradient Gram matrix: it steers selection toward the noisiest examples and makes the full batch the best solution the objective can reach. The fix costs nothing. For each example, the other candidates form an independent sample of the d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית