יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

K-Bench: בנק אקדמי לבדיקת דגמי שפה גדולים בשיחות רפואיות

K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations
בנק אקדמי שמבדיק דגמי שפה גדולים בשיחות רפואיות סכניות. הבנק כולל 125 דגמים שונים של LLMs, כולל GPT-4o, ומספק תוצאות עם רמת נכונות של 94.2%. הבנק נועד לסייע בפיתוח דגמי שפה גדולים שיכולים לספק תמיכה רפואית בטוחה.
תקציר מקורי באנגליתarXiv:2609.15855v1 Announce Type: cross Abstract: % !TEX root = ../main.tex People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leadi
קרא במקור המקורי