יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ACTR: התאמה של מחשבות ותגובות לבטיחות מרוב-לשוני במודלי LLM

ACTR: Aligning Thoughts and Responses for Multilingual Safety in Reasoning LLMs
מודל LLM של רישום והתאמה של מחשבות ותגובות לבטיחות מרוב-לשוני. המאמר עוסק בהצגת ACTR, פרקטיקה שמשפרת את ההתאמה של בטיחות מרוב-לשוני במודלי LLM. המאמר כולל דוגמאות עם תוכן לא בטוח.
תקציר מקורי באנגליתarXiv:2609.37054v2 Announce Type: replace Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual sa
קרא במקור המקורי