יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

EvalSafetyGap: פרקטיקה חיברידית לבדיקת סיכונים ב-LLM

EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures
במאמר זה, המחברים חוקרים את הבעייתיות שבבדיקת תפקודם של LLM. הם מציגים פרקטיקה חיברידית לבדיקת סיכונים ב-LLM, הכוללת סקירה משולבת, תאוריה קונצפטואלית ובדיקה מבוקרת. המחברים טוענים כי הבדיקה הנוכחית של LLM אינה מספקת וכי יש צורך בשיפור בבדיקת תפקודם של LLM.
תקציר מקורי באנגליתarXiv:2606.30219v3 Announce Type: replace-cross Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work. We introduce EvalSafetyGap as an o
קרא במקור המקורי