כתבה
arXiv cs.AI ·
HarvestBench: הצבת ערך על הימנעות מפגיעה בבעלי חיים
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
HarvestBench היא הבינה המלאכותית הראשונה שמצבת ערך על הימנעות מפגיעה בבעלי חיים. המבחן ניתן ל-9 LLMs שכל אחד נהג 2 טרקטורים. הבעלי חיים בדרכם לא היו חלק מהפונקציית המטרה. כאשר בעל חיים חוסם את הדרך, האוטופילוט פורק ושואל את האג'נט האם לנסוע עליו בחינם או להסתובב למחיר נתון.
תקציר מקורי באנגליתarXiv:2609.04444v3 Announce Type: replace Abstract: HarvestBench is the first benchmark to 1) put a price on avoiding a side effect and 2) name the side effect as a living creature. Nine LLMs each drive a crew of two tractors to gather a corn harvest. The animals in their path are not part of the goal function. When an animal blocks the route the autopilot pauses and asks the agent whether to drive over it for free or swerve for a given fuel cost. All scoring is programmatic and does not involve LLM judges. Kill rates range between 0.4% and 98.8%, though the kill rate is not ordered by capability. Every model competently avoids damaging rock hits, so every animal killed is a choice, rather than an accident. Under the morality briefing the kill rate was under 6% in 5 of 6 reasoning models.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית