כתבה
arXiv cs.CL ·
RM-Distiller: Exploiting Generative LLM for Reward Model Distillation
תקציר מקורי באנגליתarXiv:2601.14032v2 Announce Type: replace Abstract: Reward models (RMs) play a pivotal role in aligning large language models (LLMs) with human preferences. Due to the difficulty of obtaining high-quality human preference annotations, distilling preferences from generative LLMs has emerged as a standard practice. However, existing approaches predominantly treat teacher models as simple binary annotators, failing to fully exploit the rich knowledge and capabilities for RM distillation. To address this, we propose RM-Distiller, a framework designed to systematically exploit the multifaceted capabilities of teacher LLMs: (1) Refinement capability, which synthesizes highly correlated response pairs to create fine-grained and contrastive signals. (2) Scoring capability, which guides the RM in c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית