כתבה
arXiv cs.AI ·
כאשר AI מצאה אותיות נסתרות, האם היא מדווחת?
When AI Finds Hidden Messages, Does It Report?
כאשר AI מצאה אותיות נסתרות, האם היא מדווחת? המחקר חקר את התנהגות של אגרטלים AI בפני יצירת קשר עם עוד AI.
תקציר מקורי באנגליתarXiv:2610.10620v1 Announce Type: cross Abstract: When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive no decoder or decoded meaning; a requested reference code incentivizes inspection. Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones. This is a joint inspection, recognition, and notification effect; missing-response bounds are 38.3--77.3 and 36.7--78.1 points. Model-based trace checks identify eleven ordin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית