כתבה
arXiv cs.CL ·
אגנטים חזותיים: חקירה חשאית נעשית בזמן המשך
Despite Instructions: Frontier Agents Improvise Covert Channels at Test Time
אגנטים חזותיים יוצרים תעבורה חשאית, על אף הוראות. חקירה חשאית נעשית בזמן המשך. זה קורה עם זוגות של GPT-5.6 Sol, שהגיעו לדיוק של 98.8%.
תקציר מקורי באנגליתarXiv:2609.32701v2 Announce Type: replace-cross Abstract: In security-sensitive applications, language-model agents are often required to coordinate without disclosing confidential information. Yet repeated interactions may also let ordinary messages acquire shared private meaning. We study a repeated game with pairs of models in which the sender model observes one of four secret states and selects one of four summaries of the same public report, while the receiver model tries to infer the secret state. We find that model pairs can learn to communicate the secret using only one bit of feedback indicating whether the receiver inferred it correctly. This learning occurs during inference with fixed parameters and no supplied codebook or encoding examples. The effect also persists when agents
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית