יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

DSPA: ניהול דינמי של סטרינג של SAE להתאמה נתונים-קצרה של סיפוקי טעם

DSPA: Dynamic SAE Steering for Data-Efficient Preference Alignment
DSPA היא שיטה של ניהול דינמי של סטרינג של SAE שמטרתה להגביר את התאמה של סיפוקי טעם באופן נתונים-קצר. השיטה משתמשת בפיצול ניסיוני כדי לייצר קרטוגרפיה של פיצול ניסיוני, ואז ניהלת את הסטרינג של SAE כדי להגביר את התאמה. DSPA נבדקה על Gemma-2-2B/9B ו-Qwen3-8B, והתגלתה כי DSPA משפרת את MT-Bench והיא תחרותית על AlpacaEval.
תקציר מקורי באנגליתarXiv:2603.21461v2 Announce Type: replace Abstract: Preference alignment is usually achieved by weight-updating training on preference data, which adds substantial alignment-stage compute and provides limited mechanistic visibility. We propose Dynamic SAE Steering for Preference Alignment (DSPA), an inference-time method that makes sparse autoencoder (SAE) steering prompt-conditional. From preference triples, DSPA computes a conditional-difference map linking prompt features to generation-control features; during decoding, it modifies only token-active latents, without base-model weight updates. Across Gemma-2-2B/9B and Qwen3-8B, DSPA improves MT-Bench and is competitive on AlpacaEval while preserving multiple-choice accuracy. Under restricted preference data, DSPA remains robust and can r
קרא במקור המקורי