יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מעקב אחר סיכונים במערכות LLM

Stateful Guardrails for Multi-Turn LLM Systems: A Conversational Risk Accumulation Framework
חוקרים פיתחו שיטה למעקב אחר סיכונים במערכות LLM. השיטה עוקבת אחר שינויים במהלך השיחה ומזהה סיכונים פוטנציאליים. המחקר מציג גם מערכת לבחינת היעילות של השיטה.
תקציר מקורי באנגליתarXiv:2607.19361v1 Announce Type: new Abstract: Most safety guardrails for large language models (LLMs) evaluate each prompt-response pair in isolation, which misses failures that arise only over a dialogue as benign turns compose into harm. We term this Conversational Risk Accumulation (CRA): gradual intent drift, fragmented assembly of prohibited instructions, and sensitivity build-up from repeated disclosures. We propose a session-layer CRA Framework that tracks three trajectory signals: semantic drift from a session anchor, a sensitivity-weighted information accumulation graph over extracted entities, and a compliance-gradient signal capturing increasing willingness to comply. For scoring, we provide (i) an unsupervised convex fusion for attribution and ablations, and (ii) CRA-Net DA,
קרא במקור המקורי