כתבה
arXiv cs.LG ·
בדיקת תקינות של מודלי System-1 על סדרות ניתוח ביובטריולוגי
Auditing System-1 Models on Biosecurity-Relevant Benchmarks: Calibration, Selective Prediction, and Permutation Instability in a Non-Generative Model
בדיקת תקינות של מודלי System-1 על סדרות ניתוח ביובטריולוגי. המאמר עוסק בבדיקת תקינות של מודלי System-1 על סדרות ניתוח ביובטריולוגי. המחברים בדקו את המודלים על 6,020 חידושי סולם ומצאו כי המודלים נותנים תוצאות טובות על חלק מהחידושים, אך נותנים תוצאות גרועות על חלק אחר. המחברים גם גילו כי המודלים נותנים תוצאות שונות כאשר הם נותנים את אותן החידושים בסדר שונה.
תקציר מקורי באנגליתarXiv:2609.30454v1 Announce Type: new Abstract: Non-generative "System-1" models return structured probabilistic decisions in a single forward pass, without autoregressive decoding, at a small fraction of the inference cost of a generative model. This makes them of interest as inexpensive components in larger pipelines, but their reliability on biosecurity-relevant tasks has not been systematically examined. We audit one commercial System-1 model on 6,020 multiple-choice items drawn from the Weapons of Mass Destruction Proxy (WMDP), a paraphrase-robust WMDP-Bio variant, and six LAB-Bench subtasks, measuring accuracy, calibration, error detection, selective prediction, and sensitivity to the order in which answer options are presented. Accuracy is strongly task-dependent. Once the vendor's
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית