יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אל תחיל כל פעם: סטרטיפיידד אינוקולציה פומפטינג מצמצם טריג'רס של פתחי חוליה ושומר תכונות רצויות

Don't Inoculate Everything: Stratified Inoculation Prompting Narrows Backdoor Triggers and Preserves Desired Traits
סטרטיפיידד אינוקולציה פומפטינג מצמצם את הסיכוי להופעת תכונות לא רצויות ושומר תכונות רצויות. ניתן להרחיב את המודל כדי למגר את התכונות הלא רצויות גם תחת פומפטים שמבקשים אותן.
תקציר מקורי באנגליתarXiv:2609.35356v2 Announce Type: replace Abstract: Supervised fine-tuning can teach language models undesired behaviours alongside desired ones. Inoculation prompting (IP) aims to limit unwanted generalisation by requesting the undesired behaviour during training and removing the request at inference. However, undesired behaviour can still appear under unrelated prompts. IP can also hinder learning of the desired behaviour. We address these limitations in settings where both behaviours co-occur in most training examples, so filtering out examples with undesired behaviour leaves only a small clean subset. We introduce stratified inoculation prompting (SIP). SIP leverages a small clean subset to demonstrate that desired behaviour should persist without the undesired one across different con
קרא במקור המקורי