כתבה
arXiv cs.CL ·
Routing Subspaces: בדיקת פער בין הערכה לפריסה במודלים שפה מעודכנים
Routing Subspaces: Auditing Evaluation-to-Deployment Mismatch in Fine-Tuned Language Models
חוקרים פיתחו שיטה לבדיקת פערים בין הערכה לפריסה של מודלים שפה מעודכנים. השיטה מזהה הבדלים בהתנהגות המודל במהלך הערכה ופריסה.
תקציר מקורי באנגליתarXiv:2607.20436v1 Announce Type: new Abstract: Safety evaluations often assume that behavior observed during testing reflects behavior in ordinary use, but fine-tuning can break this assumption. A checkpoint can appear fixed under evaluation-style prompts while the same behavior persists under ordinary-use prompts. Output scores reveal this mismatch but do not locate it. We investigate whether the distinction is encoded in a stable internal site and introduce an approach that fits a paired activation contrast at a path-patching-informed mid-depth window, then modifies the resulting coordinate on held-out prompts. The intervention closes the evaluation-to-deployment gap in ten of twelve model--behavior settings (six of the eight settings with $n{\geq}120$ paired questions) across four full
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית