יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

איך דגלי ידע מארגנים ומבנים ידע מוסרי

How Language Models Organize and Structure Moral Knowledge
דגלי ידע מארגנים ומבנים ידע מוסרי באופן גאומטרי. המחקר חקר את האופן שבו דגלי ידע גדולים (LLMs) מארגנים ידע מוסרי. התוצאות הראו שהדגלים חושבים את היחסים בין הבסיסים המוסריים גאומטרית.
תקציר מקורי באנגליתarXiv:2608.27402v2 Announce Type: replace-cross Abstract: How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive c
קרא במקור המקורי