כתבה
arXiv cs.CL ·
אבטחה אך רגישות לעיצוב: אי-ודאות כלי בחינות LLM
Reliable but Design-Sensitive: Instrument Uncertainty in LLM Annotation
חוקרים בדקו את האמינות של מודלים LLM בחינות. התוצאות הראו שאף על פי שהמודלים נותנים תוויות אמינות בעיצוב אחד, הם משנים את התוויות כאשר החוקרים משנים את העיצוב. המחקר הראה שהשינויים בעיצוב ובבחירת המודל גרמו לשונות גדולה בתוצאות.
תקציר מקורי באנגליתarXiv:2609.35824v1 Announce Type: new Abstract: Large language models (LLMs) can give reliable labels under one setup yet change those labels when researchers make other reasonable design choices. We tested seven LLMs, 12 task designs, three independent runs, and 3,000 tweets labeled for offensive language and hate speech. Repeating the same model and task design produced high agreement (median Fleiss' $\kappa = 0.91$). Agreement fell when we changed the task design for the same tweets (median Cohen's $\kappa = 0.76$). Task design and model choice increased the variance of estimated prevalence by factors of 76.7 for offensive language and 110.6 for hate speech compared with sampling variance alone. Variation across LLM task designs reached 560-572 basis points, compared with 270-331 basis
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית