יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

PhenoBench: מפתה מהאוכלוסייה האנושית המורכבת

PhenoBench: Mapping What a Deeply Phenotyped Human Cohort Can Tell Us
PhenoBench הוא תקן לבדיקת מודלי AI על תכונות אנושיות. התקן כולל 90 משימות קליניות, ובודק את יכולתם של מודלי טבלאות ומודלי שפה להערכת תכונות אנושיות. PhenoBench יכול לשמש ככלי לבדיקת יכולות של מודלי AI, ולסייע בפיתוחם של מודלים חדשים.
תקציר מקורי באנגליתarXiv:2609.06080v1 Announce Type: new Abstract: Deeply phenotyped cohorts combine clinical, imaging, molecular, and wearable observations across timescales from seconds to years. This breadth can reveal which measurements inform which health-related questions, but heterogeneous analyses are not directly comparable. We present PhenoBench, an executable benchmark built around the Human Phenotype Project, in which more than 13,000 participants have completed the initial visit. Each question fixes the target, eligible population, timing, and allowed information; its evaluation contract specifies the split, metric, baseline, and claim boundary. The benchmark defines 90 clinically grounded tasks across 15 domains and 26 input modalities. Measurements showed question- and representation-dependent
קרא במקור המקורי