כתבה
arXiv cs.CL ·
בוטים של AI מוצאים את מה שמומחים היו?
Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions
בוטים של LLM משיגים מחקרים קליניים. ChatGPT, Claude ו-Gemini נבדקו עם 20 שאלות קליניות. התוצאות הראו שהבוטים מצליחים לשחזר 39.2% מהמחקרים הרלוונטיים.
תקציר מקורי באנגליתarXiv:2608.13786v2 Announce Type: replace-cross Abstract: Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant studies, yet the quality of retrieved evidence and factors influencing study selection remain unclear. We evaluated three general-purpose LLM chatbots (Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5) using 20 clinical questions adapted from 2026 Cochrane reviews. We simulated patient, clinician, and evidence-synthesis researcher roles and obtained four independent responses for each chatbot-role-question combination, yielding 720 responses (3 chatbots $\times$ 3 user roles $\times$ 4 repetitions $\times$ 20 review questions). Chatbots were asked to support their answers with primary clinical citations, which were b
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית