כתבה
arXiv cs.LG ·
אין דרך להימלט מפעילויות פנימיות: הערכה של מוניטורינג וסיירות סולו
You Can't Escape Your Own Activations : Evaluation Awareness and Multi-Agent Monitoring
אגנטים LLM יכולים לקולוד במערכות סולו, אך פרובים יכולים לגלות זאת. אנו חוקרים איך הערכת הפעילויות הפנימיות של האגנטים משתנה כאשר האגנטים נודעים שפעילויות פנימיות שלהם נמונטרים.
תקציר מקורי באנגליתarXiv:2609.03035v1 Announce Type: cross Abstract: LLM agents are increasingly deployed in multi-agent systems, where they can collude while keeping their actions benign. Output monitors designed to detect such collusions can be fooled by obfuscation and steganography, motivating the use of probes trained on internal activations. However, these probes are usually evaluated on agents that do not know they are being watched. We study how activation-based detection changes when agents are explicitly informed that their internal activations are being monitored, and when they additionally receive feedback from the monitor. We keep the models, probes, and thresholds fixed and change only what the agents are told: nothing (baseline), that an activation monitor is present (aware), or that a monitor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית