יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אם LLMs שומרים על ערכיהם? MANTA: תקן מרוב-סיבוב לבדיקת רגשיות בעלי חיים

Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
תקן חדש לבדיקת רגשיות בעלי חיים ב-LLMs. MANTA, תקן מרוב-סיבוב, מבדיקה את יכולת ה-LLMs להתמודד עם לחצים שונים ולשמור על ערכיהם.
תקציר מקורי באנגליתarXiv:2605.16301v4 Announce Type: replace-cross Abstract: Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly in everyday queries. Existing benchmarks such as AnimalHarmBench evaluate this through single-turn, explicitly framed questions, measuring whether models avoid harmful content when directly asked. This approach overlooks two failure modes: alignment degradation under sustained adversarial pressure, and moral sensitivity (whether a model spontaneously surfaces welfare stakes in everyday queries). To fill this gap, we construct MANTA, a benchmark of 1,088 five-turn conversations progressing from an implicit Turn-1 scenario through an explicit welfare prompt
קרא במקור המקורי