Emergent Introspective Awareness in Large Language Models
- View PDF HTML (experimental) Abstract:We investigate whether large language models can introspect on their internal states.
- It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations.
- Here, we address this challenge by injecting representations of known concepts into a model's activations, and measuring the influence of these manipulations on the model's self-reported states.
Unverified
- View PDF HTML (experimental) Abstract:We investigate whether large language models can introspect on their internal states.
- It is difficult to answer this question through conversation alone, as genuine introspection cannot be distinguished from confabulations.
- Here, we address this challenge by injecting representations of known concepts into a model's activations, and measuring the influence of these manipulations on the model's self-reported states.
Sources: Arxiv