יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מועמדות Hard-Gate בחבילת אימות מופצת

Hard-Gate Candidacy in a Deployed Validator Suite
חוקרים בדקו 13 אימותים בחבילת אימות מופצת לסוכנים יוצרים. הם מצאו שרק שני אימותים עברו בדיקת השוואה מרובה, ותשעה אחרים לא היו מובחנים מאפס. המחקר מראה שהרץ עצמו אינו אקראי ביחס לתכונה המאומתת.
תקציר מקורי באנגליתarXiv:2609.39037v1 Announce Type: cross Abstract: Before a validator can be promoted to a hard gate on a deployment pipeline, it has to be shown that its firing separates outputs that reach users in working order from those that do not. We run that screen on 13 validators in a deployed generative agent, against 550 runtime and 350 static builds labelled by downstream outcome, and report each check's marginal separation $J=\mathrm{TPR}-\mathrm{FPR}$ with Newcombe intervals and Fisher exact tests. Two checks survive correction for multiple comparisons, two more are nominal only, and the remaining nine are not distinguishable from zero, three of them because they never fired on any sampled build. Execution itself is not random with respect to the property being gated, and this replicates: acr
קרא במקור המקורי