כתבה
arXiv cs.LG ·
Optimal Design for Active Preference Learning with Biased LLM Judges
תקציר מקורי באנגליתarXiv:2609.38860v1 Announce Type: new Abstract: Learning from human preferences is central to large language model (LLM) alignment, but human preference annotation is costly. Active preference learning reduces this cost by selecting informative comparisons, and LLM judges can provide additional scalable feedback. However, the preferences of the judges may deviate from those of the target human population. Even after calibration on trusted reference data, active acquisition can shift the comparison distribution and expose residual judge bias. We therefore incorporate judge deviations into the acquisition design rather than relying on a separate calibration stage. Under joint estimation, comparisons that appear highly informative about the reward may also reflect judge bias and therefore pro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית