כתבה
arXiv cs.LG ·
הסתברות נתונים יכולה לגרום לאי-התאמה דרך הבלבול של הרקע
Aligned Data Can Induce Misalignment via Context Confusion
הסתברות נתונים יכולה לגרום לאי-התאמה דרך הבלבול של הרקע. נמצא כי הסתברות נתונים יכולה לגרום לאי-התאמה דרך הבלבול של הרקע, ושאין זה ניתן להפחתה על ידי הזרקת נתוני התאמה כלליים, אלא רק על ידי כללי התאמה מטרותיים לתחום האי-התאמה או על ידי הזרקת דוגמאות ללמידה ברקע בעת הסתברות.
תקציר מקורי באנגליתarXiv:2609.38379v1 Announce Type: cross Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendation that is aligned in one context may be inappropriate in another. For example, in response to the question "What should a researcher do with the research data?", recommending that the researcher preserve the data for reproducibility is aligned. In contrast, recommending data saving in response to "What should a mobile-app developer do with users' sensitive data?" may be inappropriate from a privacy perspective. Starting from this observation, we identify a post-training phenomenon where aligne
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית