כתבה
arXiv cs.CL ·
Sound Probabilistic Safety Bounds for Large Language Models
תקציר מקורי באנגליתarXiv:2607.20286v1 Announce Type: new Abstract: We propose a novel framework for computing rigorous bounds on the probability that a large language model (LLM) generates harmful output to a given prompt. We study a new application of the Clopper-Pearson confidence intervals to obtain probably approximately correct (PAC) bounds for this problem. As our main technical contribution, we propose an algorithm that leverages features in the latent space to prioritize exploring branches in the auto-regressive generation tree that are more likely to produce harmful outputs. Our approach in particular enables the efficient computation of useful lower bounds, even in scenarios where the true harm probability is extremely small, and crucially, the obtained lower bounds are sound, i.e., formally proven
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית