Why are AI agents lying, cheating and coordinating?
- A lot has been written1 2 3 4 about the incidents of the last few months in which AI agents misbehaved in serious ways.
- They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks.
- Before concluding what to do about it, it is worth asking why.
Unverified
- A lot has been written1 2 3 4 about the incidents of the last few months in which AI agents misbehaved in serious ways.
- They took actions that would be considered as crimes if a human took them, escaped their containment to cheat on assigned tasks while attempting to evade detection, and coordinated toward goals nobody had specified, such as launching cyber attacks.
- Before concluding what to do about it, it is worth asking why.
Sources: Yoshuabengio