כתבה
arXiv cs.AI ·
D-Score: ספקטרלי Hidden-State Signal לזיהוי Hallucination ב Large Language Models
D-Score: A Spectral Hidden-State Signal for Hallucination Detection in Large Language Models
נחשף D-Score, ספקטרלי Hidden-State Signal לזיהוי Hallucination ב Large Language Models. ה-D-Score מוצא כיצד רפרטור ההסתרה נפולט כאשר יש סתירה בין הטקסט למידע זמין במצב הפנימי של המודל. ה-D-Score מציע פתרון חדשני לזיהוי Hallucination, תוך שימוש בספקטרלי Hidden-State Signal.
תקציר מקורי באנגליתarXiv:2607.24586v1 Announce Type: cross Abstract: Large Language Models can produce fluent text that is false, unsupported by the available evidence, or inconsistent with information that appears to be internally represented by the model. We study hallucination detection from the geometry of hidden activations and introduce the D-Score, a simple spectral statistic computed from a single forward pass. For a fixed model, layer, and tolerance parameter, the D-Score counts how many singular directions of the hidden activation matrix have singular values that remain close to the leading one. We use this quantity as a hallucination score, classifying an input text as hallucinated when its D-Score is larger than a pre-defined quantity. The motivation is that, when a model processes a text that co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית