כתבה
arXiv cs.AI ·
זיהוי דרך ראשית לפתרון סתירות ידע באמצעות הנחיה באמצעות SAE
Key Path Identification for Resolving Knowledge Conflicts via SAE-based Steering
אורחות חדשות להנחיה של LLMs על ידי זיהוי דרך ראשית, שמטרתה לפתור סתירות ידע. השיטה החדשה, KPI, מציעה דרך ראשית שבאמצעותה ניתן להנחות LLMs להיות יותר נאמנים לידע הקונטקסטואלי. המחקר המדעי נערך במסגרת RAG והתוצאות היו טובות, עם עלייה של 18% בדיוק יחסית לשיטה הקודמת.
תקציר מקורי באנגליתarXiv:2609.08173v1 Announce Type: new Abstract: Sparse autoencoder (SAE)-based steering has been widely used to address knowledge conflicts by guiding LLMs to be more faithful to the contextual knowledge. Existing methods usually perform mass steering, which modifies a large batch of SAE features identified via correlation-based methods. However, due to the inaccurate correlation and the neglected feature interactions, mass steering methods fail to precisely identify the features that play the key roles in steering and introduce a large number of redundant ones, which add noise and weaken the steering effects. Our empirical studies reveal that steering only a small subset of the identified features can achieve comparable or even better performance. Motivated by this finding, we propose Key
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית