יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

תזוזת תקציב פעיל: סיווג ותיקון תוקף של תעבורת קצרה של אוטואנקודר

Active Budget Can Kill Sensitivity: Diagnosing and Repairing TopK Sparse Autoencoder Reliability
במאמר זה, חוקרים חוקרים את תקפות תעבורת קצרה של אוטואנקודר ומציעים פתרון לבעיית הפסיביות של תכונות נדירות. הם מציעים פתרון חדש, Pairwise Rank Stabilization, שמשפר את תקפות תכונות נדירות ב-8.83%.
תקציר מקורי באנגליתarXiv:2609.37857v2 Announce Type: replace Abstract: Sparse autoencoders (SAEs) are increasingly scaled to wider dictionaries to recover fine-grained structure from large language model activations. However, a feature is useful for interpretation only if it remains a stable unit of analysis when the same meaning is expressed in different surface forms. We study this reliability question for TopK SAEs via feature sensitivity. Experiments demonstrate that scaling selectively reduces the sensitivity of rare features, while common features remain comparatively stable. A controlled width\(\times k\) factorial experiment identifies the active budget k as the root cause: the degradation arises from the selection boundary rather than dictionary width alone. We attribute this failure to the geometry
קרא במקור המקורי