כתבה
arXiv cs.LG ·
התרכזות רעיונית להתערבות אמינה
Concept Concentration for Faithful Representation Intervention
במחקר זה, הצוות הציג את COCA, שיטה להתרכזות רעיונית, המאפשרת התערבות אמינה במודלי לשון. COCA מפשטת את הגבול ההחלטת בין רעיונות מסוכנים לבניים, ומאפשרת התערבות יעילה יותר.
תקציר מקורי באנגליתarXiv:2505.18672v2 Announce Type: replace Abstract: Representation intervention aims to localize and modify the representations that encode the underlying concepts in large language models (LLMs) to elicit the aligned and expected behaviors. Despite the empirical success, it has never been examined whether one could localize the faithful concepts for intervention. In this work, we explore the question in safety alignment. If the interventions are faithful, the intervened LLMs should erase the harmful concepts and be robust to both in-distribution adversarial prompts and the out-of-distribution (OOD) jailbreaks. While it is feasible to erase harmful concepts without degrading the benign utility of LLMs in linear settings, we show that it is infeasible in the general non-linear setting. To t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית