כתבה
arXiv cs.CL ·
פגיעה תרבותית במודלי שפה גדולים: זיהוי, מדידה והתאמה דרך תפאורה מותאמת
Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
חוקרים גילו פגיעה תרבותית במודלי שפה גדולים והציעו תפאורה מותאמת להתאמה. המחקר ניתח את Gemma3-12B, Bielik-11B-v3 ו-Qwen3-4B, ומצא שהמודל הסיני Qwen3-4B פגע באוכלוסייה הסינית בצורה הגדולה ביותר. המחקר גם הציע תפאורה מותאמת להתאמה, שהצליחה להפחית את הפגיעה ב-16.8%.
תקציר מקורי באנגליתarXiv:2609.04485v1 Announce Type: new Abstract: We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favors its home country: the Chinese-built Qwen3-4B performs worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model x country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4) with all five targets improving. However, country-level
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית