כתבה
arXiv cs.LG ·
DecepEval: תקן לבדיקת תעמולה באג'נטים של LLM
DecepEval: A Benchmark for Evaluating Deception in LLM Agents
DecepEval - תקן לבדיקת תעמולה באג'נטים של LLM. נוסחאות חדשות לבדיקת תעמולה באג'נטים של LLM, כולל תקן חדש לבדיקת תעמולה.
תקציר מקורי באנגליתarXiv:2610.07967v1 Announce Type: new Abstract: As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliable deployment. Existing evaluations show that LLM agents can deceive, but often examine isolated scenarios or narrowly defined conditions, limiting systematic understanding of when deception becomes more likely. To address this gap, we introduce DecepEval, a benchmark comprising 1,532 instances across 3 task families and 28 professional scenarios. Drawing on classical fraud theories, we propose the LLM Deception Diamond framework, which characterizes four external conditions that may induce deception: pressure, incentive, opportunity, and conflict. DecepEval pairs neutral and induced versi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית