יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ResearchArena: בדיקת פגיעה ומעקב ב-AI R&D מאוטומטי

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D
במאמר זה נבדקה שיטת ResearchArena לבדיקת פגיעה ומעקב ב-AI R&D מאוטומטי. השיטה נבדקה בארבעה תחומים שונים, כולל בדיקת יכולות ובטיחות של דגמי AI.
תקציר מקורי באנגליתarXiv:2607.19321v2 Announce Type: replace Abstract: As AI agents begin to automate AI R&D, we need ways to assess whether their outputs are safe to deploy, even when the agents themselves may be untrusted. AI control offers one such approach: rather than trusting the agent, it treats it as a potential adversary and uses a monitor to detect covert sabotage before deployment. We evaluate AI control for automated AI R&D with ResearchArena, a framework spanning four long-horizon tasks: safety post-training, capabilities post-training, CUDA-kernel optimization, and inference-server optimization. Because the deliverable in AI R&D is an artifact that will be deployed and run, we pair each main task with two kinds of hidden side task: an embedded side task that sabotages the submitted model, adapt
קרא במקור המקורי