כתבה
arXiv cs.CL ·
Quantifying Behavioral Tails in Black-Box Language Models
תקציר מקורי באנגליתarXiv:2609.33638v2 Announce Type: replace-cross Abstract: We introduce RareTrap, a framework for estimating the probability of severe behaviors in black box large language models (LLMs). A key challenge for probability estimation is defining a tractable distribution over the input space. To accomplish that, RareTrap uses a surrogate LLM and constructs a geometry-aware mapping from a lower-dimensional latent reference space into its token-embedding space to induce an explicit and reproducible distribution over input prompts. A response-level performance function is utilized on the response to quantify behavior severity. This enables sequential rare event simulation that concentrates evaluations on progressively more severe behaviors while preserving probability under the induced prompt dist
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית