יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מדוע הפעלה של המודל עלולה לפגוע בו? סטטוס כרייטינג לבדיקת שיטות הפעלה של LLM

Does Steering Break Your Model? A Multi-Dimensional Evaluation Suite for LLM Steering Methods
במאמר זה, המחברים פיתחו סטטוס כרייטינג לבדיקת שיטות הפעלה של LLM. הם טענו כי השיטות הקיימות לא עולות על השיטה הבסיסית של Prompt Steering באופן כללי.
תקציר מקורי באנגליתarXiv:2610.07722v1 Announce Type: new Abstract: Activation steering provides a lightweight and flexible way to control large language model (LLM) behavior. However, effective steering requires more than inducing the intended behavior: it should also limit unintended changes and remain robust across inputs and training data. Existing evaluations cover these dimensions only in fragments. As a result, the trade-offs between efficacy and side effects have not been systematically characterized. We introduce SteerScope, a two-axis, multi-dimensional evaluation suite that jointly characterizes steering outcomes and method properties through 15 metrics. We score target efficacy and side effects on language quality, task capabilities, and safety and reliability, and further assess generalization an
קרא במקור המקורי