יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אטלסי תכונות ייחוס לביקורת מנגנונית של מודלי שפה

Reference Feature Atlases for Mechanistic Auditing of Language Models
חוקרים הציעו שיטה חדשה לביקורת מודלי שפה באמצעות 'אטלסי תכונות'. השיטה מאפשרת להבין טוב יותר את המנגנונים הפנימיים של המודלים. הניסויים בוצעו על מודלים Qwen ו-Mistral.
תקציר מקורי באנגליתarXiv:2607.22570v1 Announce Type: new Abstract: Auditing a new language model usually means relearning and reinterpreting its internal features from scratch. We propose a reference feature atlas: a sparse feature library trained once on a reference panel and reused for new targets, which attach by fitting only a linear decoder. This yields two complementary views. The atlas channel reads the target on already interpreted panel features, providing a stable coordinate system across models. The residual channel learns features only from what the atlas fails to reconstruct, making "outside the reference panel" an explicit audit signal. We train leave-one-out atlases over five 7-9B instruction-tuned models and audit held-out Mistral and Qwen targets. On three controlled LoRA hidden objectives i
קרא במקור המקורי