כתבה
arXiv cs.AI ·
חקירה לקיחת קרדיט לידע במודלי שפה גדולים
Probing for Knowledge Attribution in Large Language Models
במאמר זה, המחברים חקרו את קיחת הקרדיט לידע במודלי שפה גדולים. הם פיתחו פייפלינג עצמאי שמספק נתוני אימון מותאמים אוטומטית, והציגו פרובים שמסוגלים לזהות את המקור הדומיננטי של הידע.
תקציר מקורי באנגליתarXiv:2602.22787v3 Announce Type: replace-cross Abstract: Large language model (LLM) hallucinations, meaning fluent but factually incorrect generations, fall into two types: faithfulness violations, where the model misuses provided context, and factuality violations, where answers reflect errors in internal knowledge. Proper mitigation depends on knowing which source drives each answer. We study contributive attribution, i.e. the classification of the dominant knowledge source behind each output, and show that a simple linear probe trained on hidden representations can reliably identify it. We introduce AttriWiki, a self-supervised pipeline that automatically generates labelled training data by prompting models to recall withheld entities from memory or read them from context without relyi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית