יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ACTR: התאמת מחשבות ותגובות לבטיחות מרוב-לשונית בLLMs

actr: aligning thoughts and responses for multilingual safety in reasoning llms
הצעה לשיפור בטיחות LLMs מרוב-לשונית: ACTR, פלטפורמה להתאמת מחשבות ותגובות. הצעה זו כוללת חידושים בשיפור בטיחות LLMs, כולל חישוב כוח העצבים החשוב, חישוב כוח העצבים החשוב, ואופטימיזציה של עצבים-בוחר. הצעה זו כוללת חידושים בשיפור בטיחות LLMs, כולל חישוב כוח העצבים החשוב, חישוב כוח העצבים החשוב, ואופטימיזציה של עצבים-בוחר.
תקציר מקורי באנגליתarXiv:2609.37054v1 Announce Type: new Abstract: Ensuring the safety of reasoning large language models (LLMs) across languages is essential for their reliable deployment. However, when exposed to jailbreak attacks in non-high-resource languages, these models may generate unsafe responses even when their reasoning traces identify safety risks. To address this issue, we propose aligning cross-lingual thoughts and responses (ACTR), a framework that improves multilingual safety alignment by strengthening the use of existing safety reasoning. Specifically, we first present the think gap score (TGS) to compare the normalized contributions of reasoning traces to attention outputs during response generation across languages, and use reasoning- trace substitution to measure the cross-lingual safety
קרא במקור המקורי