כתבה
arXiv cs.CL ·
חיסון אימון ביניים עם ניאולוגיזמים נלמדים
Inoculation Midtraining with Learned Neologisms
חוקרים פיתחו שיטה חדשה להפחתת סטייה במודלים גדולים של שפה. השיטה, הנקראת חיסון אימון ביניים, מלמדת את המודל לזהות התנהגות לא בטוחה ולהפריד אותה מהתנהגות תקינה. החוקרים בדקו את השיטה על מודל LLaMA ומצאו כי היא יכולה להפחית סטייה ללא פגיעה ביכולת המודל ללמוד מנתונים תקינים.
תקציר מקורי באנגליתarXiv:2609.15886v1 Announce Type: new Abstract: Large language models (LLMs) often learn both desirable and undesirable properties during post-training. We study whether midtraining, an earlier training stage, can shape which of these properties later generalise. We introduce Inoculation Midtraining, a technique that teaches a base model that unsafe behaviour belongs to a designated <quarantine_token> context, as indicated by the <quarantine_token> neologism (a new token) introduced during midtraining, and then post-trains the model on unsafe data within that context. We then evaluate the model outside the context, with the <quarantine_token> neologism excluded from the system prompt. Across supervised fine-tuning and reinforcement learning post-training regimes, we find that Inoculation M
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית