כתבה
arXiv cs.CL ·
BudgetBench: פרוטוקול ומכשיר ניסוי לבדיקת סטרטגיות זיכרון באג'נטים LLM
BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
BudgetBench הוא פרוטוקול ומכשיר שמאפשרים לבדוק סטרטגיות זיכרון באג'נטים LLM. הפרוטוקול נועד לבדוק את יעילות הזיכרון של אג'נטים LLM ולספק תוצאות עדכניות. המכשיר כולל תוכנה שמאפשרת לבדוק את הסטרטגיות ולהציג את התוצאות. BudgetBench נועד לשפר את יעילות הזיכרון של אג'נטים LLM ולספק תוצאות עדכניות.
תקציר מקורי באנגליתarXiv:2609.13149v1 Announce Type: cross Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an active-budget protocol and reference harness that treats the per-call input-token budget as the independent variable when comparing memory strategies. Holding the model, task, sampler, and decoding fixed, it sweeps budgets over 2K, 4K, 8K, 16K, and 32K tokens and records quality, budget utilization, latency, and, as a first-class outcome, budget-violation rates. The core contribution is this reusable measurement surface: a swappable MemoryStrategy contract, explicit budget enforcement, deterministic or versioned
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית