יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התאמה באמצעות הסבה נגד פרובים ללא הפסד של ניטרולביליות

Alignment via Training Against Probes Without Losing Monitorability
במאמר זה, החוקרים חקרו את השיטה של הסבה נגד פרובים, כדי למנוע התנהגות שטחית של הדגמים. הם גילו שהשיטה יעילה יותר מ-DPO ושיטות אחרות, וכי היא גם יעילה יותר במניעת פגיעה.
תקציר מקורי באנגליתarXiv:2609.38645v1 Announce Type: new Abstract: Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes that
קרא במקור המקורי