כתבה
arXiv cs.LG ·
התאמת מודלי שפה עם נתונים תצפיתיים
Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective
חוקרים את האתגרים וההזדמנויות של עדינות מודלי שפה גדולים עם נתונים תצפיתיים. הם מציגים את DeconfoundLM, שיטה להסרת השפעת גורמים מערבבים מאותות פרס. השיטה משיגה תוצאות טובות יותר מאשר שיטות אחרות.
תקציר מקורי באנגליתarXiv:2506.00152v2 Announce Type: replace Abstract: Large language models are being widely used across industries to generate text that contributes directly to key performance metrics, such as medication adherence in patient messaging and conversion rates in content generation. Pretrained models, however, often fall short when it comes to aligning with human preferences or optimizing for business objectives. As a result, fine-tuning with good-quality labeled data is essential to guide models to generate content that achieves better results. Controlled experiments, like A/B tests, can provide such data, but they are often expensive and come with significant engineering, logistical, and ethical challenges. Meanwhile, companies have access to a vast amount of historical (observational) data t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית