יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

התאמה דרך הסבה נגד פרובים ללא הפסד של ניטרולביליות

Alignment via Training Against Probes Without Losing Monitorability
במאמר זה, המחברים חקרו את האפשרות להתאמה של מודלים ללא חיסול הניטרולביליות. הם פיתחו שיטה שבה המודלים מוסבים נגד פרובים שמזהים תכונות שאינן רצויות. השיטה הצליחה להפחית את ההרסנות ולשפר את האמינות של המודלים.
תקציר מקורי באנגליתarXiv:2609.38645v1 Announce Type: cross Abstract: Models are usually aligned based on their observed outputs, using demonstrations, preference data, or reward signals. These objectives reward responses that look aligned. More capable models may learn to satisfy them without internalizing the intended behavior, for example by faking compliance during training. Such superficial compliance could be harder when the objective is defined on model internals rather than outputs. Therefore, we study probe-guided fine-tuning, using probes that detect undesired properties in model activations as a direct training signal. We evaluate linear and non-linear probes with different numbers of probes per layer across two alignment objectives: harmlessness and honesty. We find that training against probes th
קרא במקור המקורי