יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

חקירה לקיחת ידע: זיהוי מקור הידע במודלי שפה גדולים

Probing for Knowledge Attribution in Large Language Models
חקירה לקיחת ידע: זיהוי מקור הידע במודלי שפה גדולים. ניתוח של AttriWiki, פייפלין עצמאי שמספק נתוני אימון מאותו-הסוג, ומציג פרובינג של AttriWiki שמגיע ל-0.96 Macro-$F_1$.
תקציר מקורי באנגליתarXiv:2602.22787v3 Announce Type: replace Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relying on
קרא במקור המקורי