יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פרופיל נתונים נסתר של מודלי שפה גדולים

Latent Performance Profiling of Large Language Models
מאמר חדש מציע תפיסה חדשה לבדיקת מודלי שפה גדולים, על ידי פרופיל נתונים נסתר. המאמר עוסק בפיתוח של Latent Performance Profiling (LPP) - תפיסה שמספקת תיאור נסתר של תכונות המודל, כגון סובלנות לאי-וודאות וסימבוליקה.
תקציר מקורי באנגליתarXiv:2605.30018v3 Announce Type: replace-cross Abstract: Large language models (LLMs) frequently achieve impressive scores on standardized benchmarks, yet accuracy alone offers a limited view of their capabilities. Evaluating open-source LLMs on leaderboards faces persistent issues such as data contamination, a narrow task scope, and poor alignment with real-world reliability. Benchmark-based evaluations such as MMLU-Pro, BBH, or IFEval primarily capture \textit{what} a model outputs on fixed test sets, not \textit{how} it processes information, calibrates uncertainty, or structures internal knowledge. In this article, we advocate for a shift from benchmark-centric evaluation toward a complementary, \textit{state-centered intrinsic assessment} of LLMs. To this end, we introduce \textbf{La
קרא במקור המקורי