יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אימות תנאי-משטר: הערכת נכונות להתאמה ולניטור של מחליפי בטיחות

Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers
RCV מספקת תזכורת נוספת למחליפי בטיחות, על ידי הערכת נכונות של החלטותיהם. היא גם מספקת סימן תזכורת לשינויים בהפצה.
תקציר מקורי באנגליתarXiv:2608.14089v3 Announce Type: replace-cross Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. We present Regime-Conditional Verification (RCV), a lightweight wrapper that adapts an off-the-shelf safety classifier without retraining it. RCV estimates, from the classifier's internal representations, the probability that each prediction disagrees with the deployer's policy, and selectively corrects predictions likely to be wrong. The same correctness estimates also provide a label-free signal for detecting distribution shift, enabling a maintenance loop that updates the correctness est
קרא במקור המקורי