כתבה
arXiv cs.LG ·
העברת אישור לאג'נטים שאינם מאוזנים: אליגמנטציה קואליציונלית ושליטה בטוחה
Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
לצורך הבטחת בטיחות לאג'נטים AI שלאורך זמן, יש לאשר פעולות עקיבות לפני ביצוען. המאמר עוסק באליגמנטציה קואליציונלית ושליטה בטוחה.
תקציר מקורי באנגליתarXiv:2609.15803v1 Announce Type: cross Abstract: Long-running AI agents create a control problem: each action they take changes the state, which in turn affects the trajectory of future actions. If the agent is not fully aligned, then guaranteeing safety requires approving consequential actions before allowing them to be executed. But requiring human approval at every step makes attention a bottleneck. Delegating review to other AI agents raises the same alignment problem: the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guarantee that the principal fares at least as well in expectation as under a designated baseline policy. Each reviewer agent reports whether an action prop
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית