יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

תקציב פעיל יכול להרוג רגישות

Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability
חוקרים בודקים את האמינות של מקודדים דלילים מסוג TopK. הם מצאו שתקציב פעיל יכול להשפיע על רגישות התכונות. הם הציעו שיטה חדשה לשיפור היציבות.
תקציר מקורי באנגליתarXiv:2609.37857v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width\(\times k\) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry of
קרא במקור המקורי