כתבה
arXiv cs.LG ·
Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks
תקציר מקורי באנגליתarXiv:2501.04234v2 Announce Type: replace-cross Abstract: Modern artificial intelligence is supported by machine learning models (e.g., foundation models) that are pretrained on a massive data corpus and then adapted to solve a variety of downstream tasks. To summarize performance across multiple tasks, evaluation metrics are often aggregated into a summary metric, e.g., average accuracy across 10 question-answering tasks. When aggregating evaluation metrics, it is useful to incorporate uncertainty in the aggregate metric in order to gain a more realistic understanding of model performance. Our objective in this work is to demonstrate how statistical methodology can be used for quantifying uncertainty in metrics that have been aggregated across multiple tasks. The methods we emphasize are
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית