כתבה
arXiv cs.LG ·
בחינה ושיפור רובוסטנסיות של מודלי שפה גדולים לשינויי סדרה של קלט
Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
במאמר זה נחקרים ונשפרים מודלי שפה גדולים לשינויי סדרה של קלט. נצפה כי המודלים יכולים להיות רגישים להזרקות, טרוג'נים ולטיפול במטריקות אוטומטיות. נצפה כי המודלים יכולים להיות רגישים להזרקות, טרוג'נים ולטיפול במטריקות אוטומטיים.
תקציר מקורי באנגליתarXiv:2610.02432v1 Announce Type: cross Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית