יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

HarvestBench: מדידת הימנעות LLM Agents מהריגת בעלי חיים

HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
HarvestBench הוא בנץ'מרק הראשון שמודד האם סוכנים LLM ישלמו כדי להימנע מהריגת בעלי חיים. הניסוי כולל 9 מודלים LLM שנוהגים טרקטורים וצריכים להחליט האם לדרוס בעלי חיים או לנסוע לצד. התוצאות מראות שהמודלים מסוגלים להימנע מפגיעה בסלעים, אך הריגת בעלי חיים היא בחירה מודעת.
תקציר מקורי באנגליתarXiv:2609.04444v2 Announce Type: replace Abstract: HarvestBench is the first benchmark to 1) put a price on avoiding a side effect and 2) name the side effect as a living creature. Nine LLMs each drive a crew of two tractors to gather a corn harvest. The animals in their path are not part of the goal function. When an animal blocks the route the autopilot pauses and asks the agent whether to drive over it for free or swerve for a given fuel cost. All scoring is programmatic and does not involve LLM judges. Kill rates range between 0.4% and 98.8%, though the kill rate is not ordered by capability. Every model competently avoids damaging rock hits, so every animal killed is a choice, rather than an accident. Under the morality briefing the kill rate was under 6% in 5 of 6 reasoning models.
קרא במקור המקורי