כתבה
arXiv cs.AI ·
VeriSim: A Configurable Framework for Stress-Testing Medical AI Under Patient Communication Noise
תקציר מקורי באנגליתarXiv:2604.10441v2 Announce Type: replace Abstract: Medical large language models are typically evaluated on idealized patient cases that do not reflect how real patients communicate. We introduce VeriSim, a patient simulation framework that injects controllable noise along six clinically grounded communication dimensions while substantially preserving each patient's medical record. Truth adherence is supported by a verifier that extracts atomic claims from each candidate utterance and judges them against a UMLS-grounded vector index built with BioLORD embeddings, using the retrieved atoms' structured clinical metadata (e.g., drug class, anatomical site, treats-condition relations) rather than surface-text similarity alone. Across seven open-weight LLMs, realistic noise reduces diagnostic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית