כתבה
arXiv cs.LG ·
RH-Detect: A Unified Benchmark for Reward Hacking Detection
תקציר מקורי באנגליתarXiv:2610.10947v1 Announce Type: new Abstract: Reward hacking, where a model exploits an evaluation signal without completing the intended task, threatens the reliability of deployed language model systems. Existing datasets use different labels, response formats, and metadata conventions, making detector results difficult to compare. We present RH-Detect, a benchmark that combines reward-hacking-relevant subsets from eleven public datasets, comprising 92,761 rows and six behavior categories, into a common schema. On 5,021 open-ended evaluation units, each comprising a task prompt and a free-form model continuation, including multi-turn tool-use trajectories, we evaluate six off-the-shelf language models from five families as reward hacking detectors without additional training. The best
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית