יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

בייס בנוולנטי בשיחות אדם-אג'נט

Benevolent Bias in Multi-Turn Human-Agent Dialogue
בייס בנוולנטי בשיחות אדם-אג'נט יכול להיראות כאשר האג'נט מגיב בצורה חמה, אך שווה.
תקציר מקורי באנגליתarXiv:2608.29206v2 Announce Type: replace Abstract: Bias in human-agent interaction can manifest not only through hostile language but also as benevolent bias, whereby unequal treatment hides behind a warm, positive tone. To make it detectable, we operationalise benevolent bias along two dimensions, tone and treatment, yielding three classes: neutral support, overt bias, and benevolent bias. Building on these definitions, we construct BENEVDIAL, a class-balanced corpus of 362,880 multi-turn support dialogues spanning user and agent demographics, roles, and generators, to support controlled evaluation. We then test two detector families on it: off-the-shelf safety detectors and prompted large language model (LLM) judges. Our findings reveal a notable detection gap: off-the-shelf detectors r
קרא במקור המקורי