כתבה
arXiv cs.LG ·
חלוקת תקציב דינמית לבדיקה של LLM תחת תנאי משאבים קשים
Dynamic Budget Allocation for LLM Evaluation under Hard Resource Constraints
חלוקת תקציב דינמית לבדיקה של LLM תחת תנאי משאבים קשים. המאמר מציג פתרון לבעיה של קביעת גבולות זמן-לאירוע תחת תנאי תקציב קשים. הפתרון, שנקרא HARP, מאפשר לבדוק LLM תחת תנאי תקציב קשים ולקבוע גבולות זמן-לאירוע.
תקציר מקורי באנגליתarXiv:2610.07362v1 Announce Type: new Abstract: We evaluate large language models (LLMs) in multi-turn interactions through their time-to-event: the number of interaction steps required to produce an event of interest, such as a successful jailbreak or agentic task completion. Under limited compute, interactions may be terminated before the event occurs, so that event times are only partially observed (censored). Existing allocation methods for calibrating time-to-event bounds satisfy the budget only in expectation and can exceed the available budget on a particular evaluation run. Enforcing a hard constraint is particularly challenging as the cost of a trajectory is initially unknown. We introduce Hard-budget Allocation with Reflow for Predictive calibration (HARP), a budget allocation th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית