כתבה
arXiv cs.LG ·
Alignment via Training Against Probes Without Losing Monitorability
תקציר מקורי באנגליתarXiv:2609.38645v2 Announce Type: new Abstract: Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes that
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית