כתבה
arXiv cs.CL ·
אומדן העדפות יחסיות של מודלי שפה גדולים
A Consensus-Based Framework for Relative Preference Evaluation of Large Language Models
מחקר חדש מציג גישה חדשה להערכת מודלי שפה גדולים. המחקר משווה בין תשובות של מודלים שונים, כולל LLaMA, ומעריך את התשובות הטובות ביותר. הגישה החדשה מאפשרת הערכה מדויקת יותר של מודלי השפה.
תקציר מקורי באנגליתarXiv:2607.21632v1 Announce Type: new Abstract: Traditional benchmarks for LLMs primarily rely on static datasets and objective scoring metrics, which often fail to capture differences in response quality when multiple answers are acceptable. In such settings, correctness alone is insufficient to distinguish between responses that vary in clarity, completeness, and usefulness. This paper introduces a consensus-based evaluation framework that measures relative preference among model-generated responses rather than absolute correctness. Instead of evaluating outputs against a fixed ground truth, we assess how a panel of diverse LLMs ranks anonymized candidate responses to the same prompt. This approach treats aggregate inter-model agreement as a proxy for perceived response quality under bli
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית