An Empirical Study of Harness Design for Coding Agents
- To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management.
- Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 17
Unverified
- To enable component-level comparisons, we study this question with a lightweight coding harness whose execution loop is fixed while three components are varied: planning, action space, and context management.
- Across four models evaluated on SWE-Bench Verified and Terminal-Bench 2.1, we evaluate 17
Sources: Arxiv