HANDBOOK.md: A Benchmark for Long-Context Agentic Instruction Following

Global Tech Moderate confidence — 64/100

Unverified

Sources: Arxiv