כתבה
arXiv cs.CL ·
בחינה ושיפור רובוסטנסיות של מודלי שפה גדולים לשינויי סדרה של קלט
Evaluating and Improving the Robustness of Large Language Models to Input Sequence Variations
במאמר זה נחקרים ונשפרים רובוסטנסיות של מודלי שפה גדולים לשינויי סדרה של קלט. נצפה כי רובוסטנסיות זו יכולה לשפר את יכולת המודל להתמודד עם פגמים ופגמי תקינות. המאמר כולל תיאור של תכונות ופגמים של מודלי שפה גדולים, וכן תיאור של תכונות ופגמים של רובוסטנסיות של מודלי שפה גדולים.
תקציר מקורי באנגליתarXiv:2610.02432v1 Announce Type: cross Abstract: Large language models (LLMs) in production systems face prompt injections, trojans (backdoors), and manipulation of automatic quality metrics. This thesis develops models, methods, and algorithms for evaluating and improving LLM robustness to adversarial input sequence variations. We propose R_stab(f), a generative robustness metric based on the Jensen-Shannon divergence between per-step output distributions under small input perturbations. For localized attacks we prove V(h) <= 1 - R_class(h), where R_class(h) is the probability that a decision operator h keeps its decision under small perturbations. For non-localized attacks we propose a calibrated empirical model. For LLM-as-a-Judge systems we develop ASA, an adaptive evolutionary black-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית