Humans missed 1 in 3 threats approving AI agent commands across 40,000 plays
- A couple of months ago I published a small browser game: you play the human-in-the-loop for an AI coding agent, approving or denying its commands under time pressure.
- Some commands are routine (git status, npm test) and some other commands indicate your agent has been possessed and is sending your secrets to a remote server (cat ~/.aws/credentials).
- More on the threats associated with agents running commands and how to mitigate them can be found in the original post.
Unverified
- A couple of months ago I published a small browser game: you play the human-in-the-loop for an AI coding agent, approving or denying its commands under time pressure.
- Some commands are routine (git status, npm test) and some other commands indicate your agent has been possessed and is sending your secrets to a remote server (cat ~/.aws/credentials).
- More on the threats associated with agents running commands and how to mitigate them can be found in the original post.
Sources: Scalex