כתבה
arXiv cs.CL ·
HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
תקציר מקורי באנגליתarXiv:2607.25398v2 Announce Type: replace-cross Abstract: Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows. Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon. We present HANDBOOK_md, a benchmark of 65 agentic tasks modeled on how enterprise employees follow company handbooks. Each task places an agent in a self-contained company environment, a file workspace together with mock email, chat, calendar, issue-tracking, and commerce services exposed ov
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית