כתבה
arXiv cs.LG ·
Risk-Conditioned Fine-Tuning of Large Language Models
תקציר מקורי באנגליתarXiv:2609.08064v3 Announce Type: replace Abstract: Large Language Models (LLMs) are increasingly deployed in settings where rare but severe harmful generations can have significant consequences. Existing Risk-Averse RLHF addresses this issue by optimizing Conditional Value-at-Risk (CVaR), but it trains policies for fixed risk levels and therefore cannot adjust the desired degree of risk aversion at inference time. In this paper, we propose risk-conditioned RLHF, a framework that trains a single policy that provides a continuous risk-control interface, enabling users to select different degrees of risk aversion without retraining or deploying multiple risk-specific models. Experiments across multiple benchmarks demonstrate that a single risk-conditioned policy can adapt to different risk l
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית