יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שליטה מתגברת ב-LLMs באמצעות הנחיה סיבתית

Adaptive Multi-Value Control in LLMs via Causal Activation Steering
במאמר זה, המחברים מציגים פרקטיקה חדשה לשליטה ב-LLMs, המאפשרת שליטה מתגברת בערכים רבים. הם מציגים את AIMES, פרקטיקה המשתמשת בכיווני פעולה ביפולריים לערכים, ומשתמשת בקריאות ווקבולריות כצופים. הם מציגים תוצאות של AIMES ב-LangGraph וב-GPT-5, ומציעים שהפרקטיקה יכולה לשפר את השליטה ב-LLMs.
תקציר מקורי באנגליתarXiv:2609.30405v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer v
קרא במקור המקורי