Measuring Reward-Seeking by Instilling Contrastive Beliefs
- We checked this test works on models explicitly trained to favor an authority’s preferences, as well as mod
Unverified
- We checked this test works on models explicitly trained to favor an authority’s preferences, as well as mod
Sources: Openai