כתבה
arXiv cs.AI ·
MedBenchAgent: כלי לאוטומציה של בניית בסיסי ניטור רפואי
MedBenchAgent: Towards Systematic Automation of Medical VLM Benchmark Construction
MedBenchAgent: כלי לאוטומציה של בניית בסיסי ניטור רפואי. הכלי מאפשר בניית בסיסי ניטור רפואי תוך שימוש במודלי שפה גדולים ומאגרי נתונים רפואיים. הכלי כולל תיאור של תפקידי המשימה, תיאור של האקדמות, תיאור של תקן הבדיקה ותיאור של התכונות של הבסיסי הניטור.
תקציר מקורי באנגליתarXiv:2610.11312v1 Announce Type: new Abstract: Large-scale construction of medical vision-language model (VLM) benchmarks is increasingly feasible with richly annotated imaging datasets and large language models (LLMs), yet existing automation largely focuses on generating evaluation items within predefined benchmark specifications. We study the broader problem of automatically deriving the specification itself: what to evaluate, which annotations support each task, and how to translate this evidence into reliable evaluation items. We formulate benchmark construction as constrained compilation, in which the benchmark specification is progressively derived from evaluation requirements, heterogeneous annotations, and medical knowledge. Based on this formulation, we introduce MedBenchAgent,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית