יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הערכה בייסיאנית רציפה של התנהגות מודלי שפה גדולים

Sequential Bayesian Evaluation of Large Language Model Behavior
פותחים מתודולוגיה בייסיאנית להערכת התנהגותם של מודלי שפה גדולים. המחקר מדגים את הגישה החדשה בארבעה מקרי מבחן, כולל LLM עם יכולות שיחה ומודלים לחיפוש ברשת.
תקציר מקורי באנגליתarXiv:2511.10661v2 Announce Type: replace Abstract: It is increasingly important to evaluate the characteristics of systems based on large language models (LLMs). Evaluations in this context often rely on a curated benchmark set of input prompts provided to the LLM, where the output for each prompt may be assigned a binary or ordinal score and the aggregation of scores across prompts is then used as a summary evaluation. In this paper, we develop a Bayesian approach for quantifying the uncertainty that arises in such evaluation metrics as a result of the stochasticity of the LLM-based systems; the same prompt may exhibit different outcomes on repeated runs. Our framework leads naturally to a sequential evaluation, in which we leverage the Bayesian model to preferentially select which promp
קרא במקור המקורי