HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following
- View PDF HTML (experimental) Abstract:Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.
- Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon.
- We present this http URL, a benchmark of 65 agentic tasks modeled on how enterprise employees follow company han
Unverified
- View PDF HTML (experimental) Abstract:Language-model agents are increasingly deployed under standing instructions: a system prompt, a policy file, or a skills document is placed in context, and the agent is trusted to let it govern every action that follows.
- Existing benchmarks rarely test this deployment pattern directly; they measure whether an agent can complete a task, not whether a long, binding policy document actually constrains its behavior over an extended tool-use horizon.
- We present this http URL, a benchmark of 65 agentic tasks modeled on how enterprise employees follow company han
Sources: Arxiv