כתבה
arXiv cs.CL ·
JudgeMoE: אגרגציה מבוזרת ל-LLM
JudgeMoE: Distributional Aggregation for LLM-as-a-Judge
JudgeMoE הוא אלגוריתם חדש שמשפר את דיוק הציון של מודלים LLM. הוא עושה זאת על ידי אגרגציה מבוזרת של הציונים. התוצאות מראות שיפור משמעותי בדיוק הציון.
תקציר מקורי באנגליתarXiv:2610.07109v1 Announce Type: new Abstract: When an LLM judge scores an output, its score distribution retains uncertainty and disagreement information that is lost after scalar compression. We introduce JudgeMoE, a lightweight aggregator that assigns example-specific weights to cached judge score distributions and fuses them before computing a final score. A protocol study shows that score-range choice is unstable across judge--dataset settings and that soft scoring usually outperforms hard decoding. On the original 10-cell benchmark, JudgeMoE improves mean Spearman over uniform log pooling by $+0.079$. Applying the same configuration to six additional cells yields a $+0.0393$ mean gain over the strongest local single judge across 16 cells, with positive differences in 12/16 cells and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית