כתבה
arXiv cs.AI ·
Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents
תקציר מקורי באנגליתarXiv:2610.03315v1 Announce Type: new Abstract: Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equally useful for diagnosis. We present \textit{LiteTrajEval}, a lightweight architecture for budget-bounded trajectory evaluation. LiteTrajEval derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports. Evaluated on public Magentic-One-style and $\tau$-bench-style trajector
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית