יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Weight Oracles: קריאת ערכי רשת עצבית עם דגלי שפה

Weight Oracles: Reading Neural Network Weights with Language Models
מודלי שפה שהוטמעו קוראים תכונות של רשת עצבית על ידי קריאת ערכי רשת גולמיים.
תקציר מקורי באנגליתarXiv:2610.07334v1 Announce Type: new Abstract: Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this in
קרא במקור המקורי