כתבה
arXiv cs.LG ·
הפיצוח של אלגוריתמים ב-LLMs: מקרה על דגמי Markov מחבור
Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models
LLMs יכולים לצפות את הערכות הבאות מדגמי Markov מחבור דרך למידה בטקסט, אך האלגוריתם התת-יישומי נותר לא ידוע.
תקציר מקורי באנגליתarXiv:2607.22646v1 Announce Type: cross Abstract: Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations. We close this gap with a three-stage pipeline. First, we empirically compare LLM behavior against a suite of candidate algorithms and narrow the space to three classes -- though no single class explains LLM behavior across all HMM settings and sequence lengths. Second, we derive theoretical connections between the three classes and show how each can be implemented in-context by a Transformer, validating t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית