כתבה
arXiv cs.LG ·
האם LLMs יכולים לסובב את עצמם על אורכי זמן ארוכים? בחינה אמפירית של סטרטגיות תפוקה למחקר רפואי ממושך
Can LLMs Reason Over Long Horizons? An Empirical Evaluation of Context Strategies for Longitudinal Clinical Reasoning
במאמר זה נבחנה יכולתם של LLMs לסובב את עצמם על אורכי זמן ארוכים במחקר רפואי. נבחנו חמש סטרטגיות שונות לתפוקה, כולל GPT-5.
תקציר מקורי באנגליתarXiv:2610.00562v1 Announce Type: cross Abstract: Longitudinal clinical reasoning requires large language models (LLMs) to identify and integrate relevant evidence distributed across extended patient histories. Although long-context models can process increasingly large amounts of information, providing more history does not necessarily make relevant evidence more accessible or improve reasoning. We compare five context strategies (Full, Recent, Episodic, Semantic, and Hybrid) on MedLoCoMo across four open-weight LLMs, examining answer correctness, robustness to query-evidence distance, and abstention on questions with unsupported premises. Episodic and Hybrid generally achieve the strongest overall accuracy, while Recent Context degrades most as supporting evidence becomes more distant; E
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית