כתבה
arXiv cs.AI ·
בניית וביקורת טרקטוריות רציונליות קליניות לסוכני רפואה
Constructing and Evaluating Clinical Reasoning Trajectories for Medical Agent
במאמר זה נציג פרקטיקה חדשה לביקורת סוכני AI רפואיים, הכוללת ביקורת טרקטוריות רציונליות. הפרקטיקה, הנקראת MedTraj, מספקת תיאור מבוקר של טרקטוריות רציונליות, כולל תיאור של צעדי רציונליות, תיאור של ראיות, ותיאור של תוצאות. הפרקטיקה נבחנה במספר תחומים, כולל CareQA, PubMedQA, ו-CECMed.
תקציר מקורי באנגליתarXiv:2609.05090v1 Announce Type: new Abstract: Evaluation of medical artificial intelligence agents remains predominantly answer-centric, assessing only the correctness of final outputs while overlooking the quality of intermediate reasoning. In clinical settings, however, a correct answer reached through fabricated evidence or incoherent logic is as dangerous as an incorrect one. We propose MedTraj, a framework that treats reasoning trajectories as critical objects for construction, evaluation, and optimization. The pipeline generates structured multi-step reasoning chains from medical reasoning sources. Each trajectory is then parsed into clinical observations, evidence, numbered reasoning steps, and a final conclusion, and scored across five quality dimensions: coherence, evidence supp
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית