כתבה
arXiv cs.LG ·
אימות ערך כיול אחד לקופסה שחורה
Single-Query Black-Box Calibration Auditing via Logit Bias
חוקרים פיתחו שיטה לבדיקת כיול מודלי שפה גדולים. השיטה מאפשרת לבדוק את רמת הכיול של מודלים כמו LLaMA ו-GPT, אפילו כאשר הם מופעלים כ-API שחורה.
תקציר מקורי באנגליתarXiv:2609.05125v1 Announce Type: new Abstract: Evaluating the calibration of Large Language Models (LLMs) is critical for their safe deployment as zero-shot classifiers. Yet, commercial API providers increasingly hide the continuous output probabilities required by standard calibration metrics. To bypass this opacity, we demonstrate that any LLM API exposing a logit\_bias parameter can be mathematically manipulated to evaluate exact probability thresholds using strictly one query per sample. Leveraging this mechanism, we introduce a novel and provably consistent estimator of the True Calibration Error for binary tasks. Our approach therefore provides an efficient framework for auditing black-box foundation models.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית