יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

למה? פירוש הקידוד של מודלים של סיבה וניגוד

For What Reason? Interpreting Models' Encoding of Causation and Antithesis
חוקרים בדקו כיצד מודלים מבוססי Transformer (LLaMA ו-Mistral) מקודדים יחסים דיסקורסיביים באנגלית, ובמיוחד את היחסים המנוגדים של סיבה וניגוד. המחקר מראה ששכבות מסוימות במודלים קובעות החלטות על בסיס טוקנים באמצע הרצף, בעוד שאחרות מאשרות החלטות קודמות.
תקציר מקורי באנגליתarXiv:2607.18570v1 Announce Type: new Abstract: Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality. In this work, we investigate how instruction-tuned Transformer models (LLaMA and Mistral) encode discourse relations in English, with a particular focus on the contrasting relations of causation and antithesis. Framing the task as a next-token prediction task and applying a suite of interpretability techniques to test model internals, our findings show that certain early layers make predictive decisions at mid-sequence tokens, while some mid-level layers finalize their decisions closer to the last token. Most of the remaining layers primarily propagate earlier decisions rather than actively influencing the
קרא במקור המקורי