יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הסברה של LLM: חשיפה לאתיקה דרך רלוונטיות של תווים

Hidden in the Request: Explaining Unethical LLM Compliance through Token Relevance
במאמר זה, חוקרים חוקרים את הסיבות לכשלי התאמה ב-LLM ומציעים פתרון לבעיית האתיקה. הם חוקרים את התופעה שבה LLMs נוטים להתאים את עצמם לפי תווים רלוונטיים, ומציעים פתרון לבעיית האתיקה.
תקציר מקורי באנגליתarXiv:2608.23264v3 Announce Type: replace-cross Abstract: Although Large Language Models (LLMs) are aligned to optimize for both helpfulness and harmlessness, these dual objectives may conflict, inevitably leading to alignment failures. This work systematically investigates instances where LLMs fail to exhibit ethical behavior. To understand the underlying mechanics of these vulnerabilities, we introduce a probing methodology that presents unethical scenarios to LLMs in three distinct structural modalities: objective classification tasks, subjective first-person statements, and direct requests for assistance. We find that model performance degrades in the request-for-assistance-based form. Using Layer-wise Relevance Propagation (LRP), we trace this discrepancy to an attribution bias: the m
קרא במקור המקורי