כתבה
arXiv cs.AI ·
SchemeArena: בדיקת דרכים לסכמות באג'נטים LLM
SchemeArena: Factorized Stress Testing of Scheming in LLM Agents
במאמר זה, נחקרה בדיקת דרכים לסכמות באג'נטים LLM. המאמר עוסק בפיתוח של SCHEMEARENA, בנקודת ציון 400-סקנריה לבדיקת דרכים לסכמות באג'נטים LLM. המאמר גם עוסק בפיתוח של SCOUT, משקף שכן של בדיקת דרכים לסכמות באג'נטים LLM.
תקציר מקורי באנגליתarXiv:2609.08126v1 Announce Type: new Abstract: We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a small number of scenarios, limiting the ability to isolate how these conditions shape an agent's propensity or capability to scheme. This limited scale and task diversity also restrict coverage of realistic deployment settings and the range of scheming strategies that can be observed. To this end, we introduce SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית