כתבה
arXiv cs.AI ·
הבעיה של המטרה להתאמה: הפסדים מוסריים של בני אדם, מערכות AI ומעצביהן
The Alignment Target Problem: Divergent Moral Judgments of Humans, AI Systems, and Their Designers
הפרויקט של התאמת התנהגות מכונה לערכים אנושיים עוררה בעיה יסודית: מי הצפוי לקבוע את תפיסותיו המוסריות של AI. חקירות של ענפי ערכי סוג-מערכת פקפקו בהנחה שהתקן הנכון הוא כיצד יתנהגו בני אדם במצב נתון. המאמר הזה חודר עוד צעד נוסף בכך שבודק שתי אפשרויות נוספות: ראשית, שביקורתיות של התנהגות AI תשתנה כאשר מופיעות ראיות למקורה האנושי; ושנית, שאנשים ישפוטו את בני אדם שתכננו AI כשונה מכלפי של המכונות או האנשים שהם מוליכים למערכת. ניסוי עם 1,002 אמריקאים חשף פסדים מוסריים כאשר השתמשו ב-1,002 אמריקאים כדי לבחון פסדים מוסריים.
תקציר מקורי באנגליתarXiv:2604.24155v4 Announce Type: replace-cross Abstract: The project of aligning machine behavior with human values raises a basic problem: whose moral expectations should guide AI decision-making? Much alignment research assumes that the appropriate benchmark is how humans themselves would act in a given situation. Studies of agent-type value forks challenge this assumption by showing that people do not always judge humans and AI systems identically. This paper extends that challenge by examining two further possibilities: first, that evaluations of AI behavior change when its human origins are made visible; and second, that people judge the humans who program AI systems differently from either the machines or the human actors they are compared against. An experiment with 1,002 U.S. adul
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית