כתבה
arXiv cs.AI ·
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
תקציר מקורי באנגליתarXiv:2603.10044v4 Announce Type: replace Abstract: Safety benchmarks usually test "bare" models that receive prompts and output responses, but real-world deployments "wrap" those models in complex scaffolds. How much do these scaffolds affect model safety as measured by benchmarks? We test six leading models on four pre-registered safety benchmarks with a direct API and three scaffolds: ReAct, multi-agent, and map-reduce. We conducted 60,112 scored evaluations. On average, how safety is measured matters more than scaffolding does: we find that using a multiple choice vs. open-ended format for otherwise-identical benchmark items changes measured safety by about 5-20 percentage points (pp). The two formats are scored with different methods (answer extraction and an LLM judge), so the gap is
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית