כתבה
arXiv cs.AI ·
לתפוס את המעשה: פרובים מזהים הפרת אמונים ומגיעים למסקנות
Caught in the Act: Probes Effectively Detect Sabotage and Catch Unverbalized Deception
פרובים עוזרים לזהות הפרת אמונים והפרת תפקוד במודלי LLM. הם גם עוזרים לזהות דליפות והפרת סודיות. הפרובים עובדים בעזרת נתונים שנאספו ממקורות שונים, כמו טקסטים, תמונות וסרטונים. הם גם עוזרים לזהות דליפות של מידע פרטי.
תקציר מקורי באנגליתarXiv:2610.12445v1 Announce Type: cross Abstract: Recent incidents have highlighted the challenge of monitoring LLM agents and the danger of models deceiving people. We show that white-box deception detection via probes can be scaled up to frontier monitoring settings by collecting the largest deception dataset to date for training probes and introducing a novel probe architecture which can aggregate information across many layers and tokens. Our probes achieve 98.8% AUC in SHADE-Arena, surpassing an Opus 5.5 text-monitoring baseline, and show improved efficacy as the underlying model is scaled up. To push our probes to their limit, we test them on several cases where deception cannot be determined from the context alone. In these cases, which we refer to as introspective deception, the gr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית