כתבה
arXiv cs.CL ·
שיפור אמינות LLM בזמן בדיקה
A Removal Based Approach to Improve LLM Faithfulness at Test-Time
חוקרים מציעים שיטה חדשה לשיפור אמינות מודלי LLM בזמן בדיקה. השיטה כוללת הסרת מושגים שאינם מוזכרים בהסבר המודל וביצוע שאילתה חוזרת עם הקלט המופחת. השיטה משפרת את אמינות ההסברים לעומת שיטות קיימות.
תקציר מקורי באנגליתarXiv:2609.04343v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for consequential decisions, making their explanations an important tool for auditing model behavior. Unfortunately, these explanations can be unfaithful, failing to reflect the actual reasoning underlying the model's decisions. We consider a setting in which an LLM provides both an answer and an explanation in response to a question. We identify two distinct dimensions of unfaithful explanations: incompleteness, meaning that the explanation omits factors that influence the answer, and unsoundness, meaning that the explanation cites factors that did not influence the model's answer. Existing approaches to improving LLM faithfulness include training-time methods, which require access to mode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית