כתבה
arXiv cs.AI ·
אל תחילו את כולם: סטרטיפייד אינוקולציה פומפטינג צרה את הסימנים של הפתח האחורי ושומרת את התכונות הרצויות
Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
סטרטיפייד אינוקולציה פומפטינג צרה את הסימנים של הפתח האחורי ושומרת את התכונות הרצויות. המחקר מציע שיטה חדשה להגבלת התנהגות לא רצויה במודלי תקשורת.
תקציר מקורי באנגליתarXiv:2609.35356v3 Announce Type: replace Abstract: Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours co-occur in most training examples, so filtering out examples with undesired behaviour leaves only a small clean subset. We introduce stratified inoculation prompting (SIP). SIP leverages a small clean subset to demonstrate that desired behaviour should persist without the undesired one across different con
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית