יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

BenchShield: כלי רפואי להגנה על תקיפה של תגמולים בתשתיות כניסה למודלי LLM

BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
BenchShield הוא כלי רפואי שמטרתו להגן על תקיפה של תגמולים בתשתיות כניסה למודלי LLM. הוא משתמש במודל חיים סופי של אירועי תגמול-רלוונטיים כדי לזהות תקיפה של תגמולים.
תקציר מקורי באנגליתarXiv:2609.11028v1 Announce Type: cross Abstract: LM-agent benchmarks increasingly function as interactive evaluation infrastructure. Agents observe state, call tools, modify workspaces, submit artifacts, and receive rewards from outcome procedures. This interactivity makes evaluations vulnerable to reward hacking: an agent improves its measured score by exploiting the reward-relevant trajectory instead of solving the intended task. Existing defenses rely largely on task-specific patches, prompt instructions, or post-hoc detectors. They do not provide reusable evidence that a concrete run remained within its intended evaluation boundary. This paper presents BenchShield, a model-backed instrumentation layer for reward integrity in LLM-agent evaluation. BenchShield grounds detection in a fin
קרא במקור המקורי