כתבה
arXiv cs.LG ·
Asymptotic Universal Alignment: A New Alignment Framework via Test-Time Scaling
תקציר מקורי באנגליתarXiv:2601.08777v2 Announce Type: replace Abstract: Aligning large language models (LLMs) to serve users with heterogeneous and potentially conflicting preferences is a central challenge for personalized and trustworthy AI. We formalize an ideal notion of universal alignment through test-time scaling: for each prompt, the model produces $k\ge 1$ candidate responses and a user selects their preferred one. We introduce $(k,f(k))$-robust alignment, which requires the $k$-output model to have win rate $f(k)$ against any other single-output model, and asymptotic universal alignment (U-alignment), which requires $f(k)\to 1$ as $k\to\infty$. Our main result characterizes the optimal convergence rate: there exists a family of single-output policies whose $k$-sample product policies achieve U-align
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית