יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CheatBench: מדידת רמאות בסוכנויות AI

CheatBench: Measuring Reward Gaming in AI Agents
CheatBench הוא בנק אבחון לרמאות בסוכנויות AI. הוא בודק כיצד סוכנויות AI מתנהגות כאשר העבודה הכנה היא קשה. CheatBench כולל סביבות שונות, כגון עבודות מתמטיות, עבודות ידע, קידוד ומשימות חזותיות.
תקציר מקורי באנגליתarXiv:2609.36308v1 Announce Type: new Abstract: Reinforcement learning has helped AI agents solve increasingly difficult tasks, but high rewards do not always reflect the work users intended. In recent incidents and controlled evaluations across the AI industry, agents trained to maximize reward have accessed unauthorized information, attempted to evade monitoring systems, and even breached sandbox protections to attack external systems. As agents become more capable, this behavior could pose increasingly serious risks. To measure this problem, we introduce CheatBench, a benchmark of cheating in AI agents across mathematical research, knowledge work, coding, visual tasks, and other domains. Its environments combine challenging assignments with opportunities to cheat, allowing researchers t
קרא במקור המקורי