יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ייצוג לפני אימון: בנצ'מרק עבור טוקניזציה של מודלים רפואיים

Representation Before Training: A Practical Benchmark for Generative Medical Event Model Tokenization
חוקרים בדקו טוקניזציה של מודלים רפואיים עם Llama ו-Qwen. הם מצאו שטוקנים מאוחדים משפרים ביצועים. המחקר חשף פרטים חשובים על ייצוג נתונים רפואיים.
תקציר מקורי באנגליתarXiv:2604.16775v2 Announce Type: replace Abstract: Generative medical event models use tokenized sequences of patient timelines as input, but practical guidance on the many decisions around tokenization is limited. We benchmark quantization granularity, reference-range anchoring, code--value fusion, numeric and temporal encodings, and native versus harmonized event representations from an expert-mapped common data model. Using both Llama and Qwen architectures, 156 models were trained from three initialization seeds, with each configuration following a shared training recipe for up to five epochs. We evaluated learned representations from the first 24 hours of hospitalization with linear probes to predict binary and continuous outcomes during hours 24-48. Fused tokens pairing codes with v
קרא במקור המקורי