כתבה
arXiv cs.AI ·
Generative Interpretability via Scalable Neuro-Symbolic Models
תקציר מקורי באנגליתarXiv:2609.13529v1 Announce Type: cross Abstract: As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward \emph{generative interpretability}, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית