Self-generated prompt injections in compaction summaries · OpenAI Alignment
- Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable.
- Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug
Unverified
- Our conclusion was that this behavior was extremely rare, did not confer an obvious reward advantage, and was monitorable.
- Our top hypothesis is that issues around summary termination contributed to this behavior, though we have not established a causal connection, and we have addressed a related bug
Sources: Openai