יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TRACE: ייחוס מסלול ומחיקה קונטראסטיבית

TRACE: Trajectory Return Attribution and Contrastive Erasure for Multi-Turn Safety
TRACE הוא אלגוריתם חדש לשיפור בטיחות של מודלי שפה גדולים. הוא משתמש בייחוס מסלול ומחיקה קונטראסטיבית כדי לצמצם את הסיכון של תגובות מסוכנות. האלגוריתם נבדק על מודלים שונים והראה תוצאות מבטיחות.
תקציר מקורי באנגליתarXiv:2610.01323v1 Announce Type: new Abstract: Safety-aligned large language models (LLMs) often refuse a harmful request but comply once the same goal is spread over several turns. Preference objectives score whole responses to single prompts, so their training loss alone cannot control risk on unseen histories. Our analysis gives sufficient conditions under which suppression at supervised single-turn contexts yields a bound on multi-turn trajectory risk. The bound accounts for coverage, transfer slack, and leakage, and characterizes contraction relative to a base-policy risk budget evaluated on the trained policy's contexts. TRACE (Trajectory Return Attribution and Contrastive Erasure) turns this principle into a token-level objective. On the safe response, each token is weighted by the
קרא במקור המקורי