יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

גאומטריה של חוסר ידיעה: LLMs יודעים כימו להפחית ביינים Bayes

The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors
LLMs יודעים כימו להפחית ביינים Bayes. המחקר חושף כי LLMs עושים שימוש בגאומטריה של חוסר ידיעה כדי להפחית את ההשפעה של הביינים Bayes. התוצאות המחקר חשובות לפיתוח LLMs יעילים יותר.
תקציר מקורי באנגליתarXiv:2609.02959v1 Announce Type: cross Abstract: What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the model falls back on when uncertain. This structure --- which we term the \emph{direction of ignorance} --- appears in all four model families examined (\texttt{Llama}, \texttt{Qwen}, \texttt{Gemma}, and \texttt{Pythia}), ranging from 0.4B to 405B parameters. Projecting the final prediction state onto this direction yields a per-token \emph{prior loading factor} $\lambda$, which, empirically, declines steadily as the context becomes more informative. Formally, the same projection decomposes the predic
קרא במקור המקורי