כתבה
arXiv cs.CL ·
פרופיל ביצועים מוטמעים של מודלי שפה גדולים
Latent Performance Profiling of Large Language Models
Latent Performance Profiling (LPP) הוא כלי להערכת מודלי שפה גדולים. LPP מספק מדדים סקלריים לתיאור התנהגות המודל. המחקר מראה כי מודלים עם ציונים דומים יכולים להציג פרופילים מוטמעים שונים.
תקציר מקורי באנגליתarXiv:2605.30018v3 Announce Type: replace Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. Evaluating open-source LLMs on leaderboards faces persistent issues such as data contamination, a narrow task scope, and poor alignment with real-world reliability. Benchmark-based evaluations such as MMLU-Pro, BBH, or IFEval primarily capture \textit{what} a model outputs on fixed test sets, not \textit{how} it processes information, calibrates uncertainty, or structures internal knowledge. In this article, we advocate for a shift from benchmark-centric evaluation toward a complementary, \textit{state-centered intrinsic assessment} of LLMs. To this end, we introduce \textbf{Latent P
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית