יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

JevAdvBench: בנק אדוונס והתקפות שחור-קופסה ללמידת מודלים להחלטות מסוגלות

JevAdvBench: A Benchmark and Black-Box Attacks for Reinforcement Learning for Calibrated Decisions Models
בנק אדוונס והתקפות שחור-קופסה ללמידת מודלים להחלטות מסוגלות. המאמר מציג בנק אדוונס והתקפות שחור-קופסה לבדיקת רובוסטנסיות של מודלי RLCD. הבנק כולל 812 שאלות מסוגים שונים ו-9,744 גרסאות של השאלות עם שינויים קטנים.
תקציר מקורי באנגליתarXiv:2609.31142v1 Announce Type: cross Abstract: Models trained with reinforcement learning for calibrated decisions (RLCD), such as Jev, answer a typed question about an input, the state, with a probability, a choice, or a score, and software acts on the answer without a person reading it. Their robustness has not been measured: adversarial benchmarks score what a model generates or executes, whereas a typed model generates nothing and returns a well-formed answer even when manipulated. Measurement is also hard, because identical requests can return different answers, most available labels come from the model itself, and the API preprocesses each request out of view. Our key idea is to score each attacked decision against the model's own clean decision rather than against labels, and to
קרא במקור המקורי