כתבה
arXiv cs.AI ·
AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection
תקציר מקורי באנגליתarXiv:2609.05899v1 Announce Type: cross Abstract: Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit model performance. To bridge this gap, we propose AlignDiff, a preference data filtering framework driven by intrinsic model signals. AlignDiff first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information from them. AlignDiff is evaluated on two widely used model families (LLaMA and Qwen) and three benchmar
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית