כתבה
arXiv cs.AI ·
הפער הקונבנציונלי: כיוון כלפי מדידת תקשורת בלתי-מובהרת בבדיקות AI
The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation
המאמר חוקר את הפער הקונבנציונלי בבדיקות AI, ומציע דרך למדידת תקשורת בלתי-מובהרת בין AI לבין בני-אדם. המחקר משתמש במשחק הקלפים Hanabi כדי לבדוק את הצעתו.
תקציר מקורי באנגליתarXiv:2609.11489v1 Announce Type: new Abstract: Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions---shared protocols for reading meaning beyond the literal message---which AI-AI benchmarks may not capture. We propose the \emph{convention gap}, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (hanab.live), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, $-$0.7~pp in A
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית