כתבה
arXiv cs.LG ·
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.14896v1 Announce Type: cross Abstract: A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring open-ended exploration and pluralistic perspectives, such as scientific ideation and creative writing. We present MoDA (Mode-conditioned Diversity Alignment), an online post-training RL algorithm that jointly optimizes generation quality and diversity, inspired by the coordination perspective in multi-agent reinforcement learning (MARL). MoDA trains a single shared LLM policy conditioned on abstract numbered roles, where each role acts as an agent competing to produce outputs distinct from the others. This formulation
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית