יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Towards Verifiable Transformers: Solver-Checkable Circuit Explanations

תקציר מקורי באנגליתarXiv:2605.24033v2 Announce Type: replace Abstract: Mechanistic interpretability typically discovers circuits and then argues what they do from examples and ablations. We introduce Verifiable Transformers, a framework for turning task-localized circuits into bounded, solver-checkable claims: projected functional equivalence, task-relevant invariance, edge necessity, and robustness to continuous final-residual perturbations. At small scale, we directly verify all four properties for quote-closing and bracket-type circuits, including program-mediated circuits whose attention selection is entirely symbolic. At GPT-2 scale, we remove LayerNorm from a sparsemax/LeakyReLU model after training with a +0.0087 OpenWebText loss increase, replace retained attention heads with synthesized restricted-D
קרא במקור המקורי