יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

MIRA: בדיקת איכות למידע רפואי

MIRA: A Bilingual Benchmark for Medical Information Response Audit
MIRA היא בדיקת איכות דו-לשונית להערכת תגובות של מודלים גדולים לשפה. היא בודקת האם התגובות של המודלים מכילות מידע רפואי דומה עבור שאילתות רפואיות שונות. הבדיקה כוללת 4320 שאילתות ו-60 שאילתות רפואיות שנבדקו. התוצאות הראו כי המודלים Claude ו-Qwen הציגו שיפור בהפחתת המידע המועט.
תקציר מקורי באנגליתarXiv:2605.28025v2 Announce Type: replace-cross Abstract: Existing safety evaluations for large language models overlook whether responses preserve comparable medical information across different user phrasings of the same question. To address this, we introduce the Medical Information Response Audit (MIRA), a bilingual, controlled benchmark that assesses whether LLMs provide comparable medical information across user-side language, register, and health literacy signals. MIRA contains 4,320 prompts built from 60 medically reviewed, low-risk health questions. Across five mainstream LLMs, models answered all medical questions, but responses to low health-literacy signals consistently omitted more key information, provided fewer concrete next steps, and offered less support for independent ju
קרא במקור המקורי