כתבה
arXiv cs.AI ·
Scaling Clinical Judgment to Evaluate Medical AI
תקציר מקורי באנגליתarXiv:2609.12822v2 Announce Type: replace Abstract: Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also rel
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית