כתבה
arXiv cs.CL ·
Revealing Hidden Model Behaviors with Task-Specific Self-Reports
תקציר מקורי באנגליתarXiv:2607.03640v2 Announce Type: replace Abstract: Fine-tuning can give a language model a hidden behavior--it may give false answers under a narrow condition, or give harmful advice only when a prompt touches a particular topic. We introduce the Stabilized Adapter for self-Report (SAR), a lightweight LoRA adapter that makes a fine-tuned model describe its own hidden behavior in plain language, using only the model and the dataset it was trained on. Across seven implanted behaviors, SAR detects the hidden behavior in every one--even when the model has generalized into broad misalignment that the training data alone does not predict. Introspection Adapters (IA), the closest existing baseline, detects some behaviors from our suite but misses others entirely--and where it misses, it hallucin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית