כתבה
arXiv cs.LG ·
כמה מסמנים יש לכם?
How many labelers do you have? A closer look at gold-standard labels
חוקרים את התהליך של איסוף ואיחוד תוויות למודלים. המחקר מראה כי גישה למידע תווית לא מאוחד יכולה לשפר את הדיוק של המודלים. התוצאות מוצגות על סמך ניתוח סטטיסטי.
תקציר מקורי באנגליתarXiv:2206.12041v3 Announce Type: replace-cross Abstract: The construction of most supervised learning datasets revolves around collecting multiple labels for each instance, then aggregating the labels to form a type of "true" label. We question the wisdom of this pipeline by developing a (stylized) theoretical model of this process and analyzing its statistical consequences, showing how access to non-aggregated label information can make training well-calibrated models more feasible than it is with cleaned labels. The entire story, however, is subtle, and the contrasts between aggregated and fuller label information depend on the particulars of the problem, where estimators that use aggregated information exhibit robust but slower rates of convergence, while estimators that can effectivel
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית