יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

K-Bench: בנק אקדמי לבדיקת דגימות שפה גדולה בשיחות רפואיות

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
בנק אקדמי שפותח לבדיקת דגימות שפה גדולה בשיחות רפואיות סיכניות. הבנק כולל 125 דגימות של מודלי שפה ובודק את יכולתם לטפל במצבי רפואיים סיכניים.
תקציר מקורי באנגליתarXiv:2609.15855v1 Announce Type: new Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading
קרא במקור המקורי