יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אימות תנאי-משטר: הערכת נכונות להתאמה ולניטור של מחלצי בטיחות

Regime-Conditional Verification: Correctness Estimation for Adapting and Monitoring Safety Classifiers
RCV מספקת הערכת נכונות למחלצי בטיחות, כדי להתאים אותם למדיניות הרצויה ולזהות שינויים בהתפלגות. היא חושפת כלי חדשני לאימות תנאי-משטר.
תקציר מקורי באנגליתarXiv:2608.14089v3 Announce Type: replace Abstract: Safety classifiers deployed with large language models often fail for two reasons: their decisions reflect the policy learned during training rather than the deployer's desired policy, and their performance degrades as deployment traffic evolves. We present Regime-Conditional Verification (RCV), a lightweight wrapper that adapts an off-the-shelf safety classifier without retraining it. RCV estimates, from the classifier's internal representations, the probability that each prediction disagrees with the deployer's policy, and selectively corrects predictions likely to be wrong. The same correctness estimates also provide a label-free signal for detecting distribution shift, enabling a maintenance loop that updates the correctness estimatio
קרא במקור המקורי