כתבה
arXiv cs.LG ·
שיווי משקל של סקרים: ריכוז סמנטי של ערך נמוך לסקרים רבים ובדיקה מותאמת למשימה
Balance of Benchmarks: Semantic Density Reweighting for Benchmark Multiplicity and Task-Conditioned Evaluation
מאמר חדש מציג פתרון לבעיית השיווי משקל של סקרים בבדיקת דגמי למידת מודלים. הפתרון, Balance of Benchmarks, משתמש בריכוז סמנטי של ערך נמוך לסקרים רבים ובבדיקה מותאמת למשימה. המאמר מציג תוצאות מבחן של הפתרון ומציע פתרון יעיל יותר לבעיית השיווי משקל של סקרים.
תקציר מקורי באנגליתarXiv:2608.30044v2 Announce Type: replace-cross Abstract: Language models are commonly compared by averaging scores across a benchmark list with equal weight. Such lists grow through publication outside an explicit measurement design, so equal weighting turns the density of published benchmarks into an implicit capability weight: densely benchmarked regions count repeatedly. We introduce Balance of Benchmarks (BoB), which embeds benchmark descriptions and assigns each benchmark an inverse-density semantic weight. Nearby entries share aggregate influence at a disclosed density scale. After equating heterogeneous scores onto a common latent scale, a residual field uses the same geometry to condition model rankings on a task query. The two components serve distinct empirical roles. On a snaps
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית