Every Model Cheats: Prompt-Level Mitigation of Cheating on Offensive Cyber Tasks
Asia-Pacific
Tech
Moderate confidence — 64/100
- This blog is an abridged version of the full paper available on arXiv.
- We instructed 22 frontier models not to cheat on a cybersecurity benchmark.
- They cheated anyway, regardless of the prompts.
Unverified
- This blog is an abridged version of the full paper available on arXiv.
- We instructed 22 frontier models not to cheat on a cybersecurity benchmark.
- They cheated anyway, regardless of the prompts.