כתבה
arXiv cs.LG ·
כיצד כללי ציון הפניה יוצרים חיזוי LLM
How Proper Scoring Rules Shape LLM Forecasting
במאמר זה נבחן כיצד תפריטי פרס נכונים יוצרים את התנהגות חיזוי LLM. נבדקו חמש כללי ציון הפניה שונים כמטריצות אימון לחיזוי תצפיות בעולם האמיתי. המודל שאומנו על פי כללי ציון הפניה Brier הציג את הציון הנמוך ביותר של Brier ואת ה-AUC-ROC הגבוה ביותר.
תקציר מקורי באנגליתarXiv:2608.28482v2 Announce Type: replace Abstract: This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through diff
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית