כתבה
arXiv cs.AI ·
ארבעים גוונים של כחול
Forty Shades of Blue: Quality-Diversity Alignment via Mode-Conditioned Reinforcement Learning
MoDA היא אלגוריתם RL שמשפר את האיכות והגיוון של מודלים כמו Qwen. היא מאפשרת יצירת תוצאות מגוונות ואיכותיות יותר.
תקציר מקורי באנגליתarXiv:2609.14896v1 Announce Type: cross Abstract: A notable byproduct of LLM alignment training is mode collapse: the progressive loss of output diversity that narrows a model's expressivity at inference time. This degradation is especially limiting for applications requiring open-ended exploration and pluralistic perspectives, such as scientific ideation and creative writing. We present MoDA (Mode-conditioned Diversity Alignment), an online post-training RL algorithm that jointly optimizes generation quality and diversity, inspired by the coordination perspective in multi-agent reinforcement learning (MARL). MoDA trains a single shared LLM policy conditioned on abstract numbered roles, where each role acts as an agent competing to produce outputs distinct from the others. This formulation
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית