יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

BudgetBench: תקן-תקציבי לבדיקת סטרטגיות זיכרון באג'נטים של שפה גדולה

BudgetBench: A Budget-Tiered Protocol and Pilot Harness for Memory Strategy Evaluation in Local Large Language Model Agents
BudgetBench הוא תקן-תקציבי שמאפשר לבדוק סטרטגיות זיכרון באג'נטים של שפה גדולה. התקן משמש לבדוק את יעילות הזיכרון של האג'נטים ולזהות נקודות עבודה. התקן כולל תקן-תקציבי, חבילת תוכנה וכלים לבדיקת סטרטגיות זיכרון.
תקציר מקורי באנגליתarXiv:2609.13149v1 Announce Type: new Abstract: For local large language model agents, active context is a scarce resource: memory capacity, prefill latency, cache growth, and service objectives all constrain how many input tokens each call can afford. We present BudgetBench, an active-budget protocol and reference harness that treats the per-call input-token budget as the independent variable when comparing memory strategies. Holding the model, task, sampler, and decoding fixed, it sweeps budgets over 2K, 4K, 8K, 16K, and 32K tokens and records quality, budget utilization, latency, and, as a first-class outcome, budget-violation rates. The core contribution is this reusable measurement surface: a swappable MemoryStrategy contract, explicit budget enforcement, deterministic or versioned gr
קרא במקור המקורי